AGP Picks
View all

iKala upgrades TMMLU+ to test Traditional Chinese AI skills and agents

10 hours ago
By AI, Created 10:43 UTC, Sep 15, 2026, AGP -

iKala released TMMLU+ v1.1 in September 2026 to improve Taiwan’s ability to independently evaluate AI systems in Traditional Chinese. The updated benchmark expands from answer accuracy toward agent and multimodal task performance, aiming to help enterprises pick AI that fits local workflows, data and regulatory needs.

Why it matters: - Taiwan’s AI buyers need a local yardstick, not just global leaderboard scores. - TMMLU+ is designed to measure whether AI systems understand Traditional Chinese context, laws, policies and business use cases in Taiwan. - iKala says the benchmark supports sovereign AI by strengthening the country’s control over data, evaluation standards and AI selection.

What happened: - iKala advanced TMMLU+, its open-source Traditional Chinese AI benchmark, to strengthen Taiwan’s independent AI evaluation capability. - The company released TMMLU+ v1.1 in September 2026. - TMMLU+ v1.1 is available on Hugging Face and GitHub. - The benchmark began as TMMLU+ v1, a large-scale test of multitask language understanding for large language models in Traditional Chinese. - TMMLU+ v1 includes 22,690 multiple-choice questions across 66 subjects. - The benchmark covers material from elementary education through professional domains. - The research behind TMMLU+ was published at COLM 2024. - International research teams including Google and Meta have used the benchmark in evaluation work.

The details: - Sega Cheng, iKala co-founder and chairman, defines sovereign AI as four layers: compute, models, data and applications. - Cheng also argues that sovereign AI must include the ability to judge whether AI systems meet local needs. - iKala made TMMLU+ open source so industry and academia could use a shared standard. - iKala says Chinese language ability alone does not guarantee understanding of Taiwan’s local laws, public policies, cultural references, professional terms or daily context. - The company sees trustworthy Traditional Chinese datasets and evaluation frameworks as essential as AI becomes a primary interface for search, decision-making and information access. - Gartner predicts that by 2030, up to 90% of generative AI solutions will use domain-specific models. - That shift moves enterprise decision-making from finding any model to finding the right model or system for a specific business scenario. - Public benchmark scores may not predict how a model performs in real workflows. - Enterprises can invest heavily in proof-of-concept projects and still find that a model underperforms in Traditional Chinese business settings or fails local requirements. - iKala is extending its evaluation work from LLM testing to agent evaluation. - The next stage of testing will assess whether AI can understand intent, use tools correctly, process text, images and audio, and complete end-to-end enterprise tasks. - iKala says the most suitable AI may be the system that performs best on an organization’s own data, workflows, requirements and cost constraints rather than the one with the highest global score. - TMMLU+ v1.1 includes a systematic review of all 66 subject areas. - The review gave special attention to laws, regulations, public policies and other time-sensitive content. - Invalid, incomplete or corrupted questions were corrected or removed. - Questions with multiple plausible answers or ambiguous choices underwent additional review by domain experts. - Questions without a single defensible answer were excluded.

Between the lines: - iKala is positioning evaluation infrastructure as strategic AI infrastructure, not just a technical side project. - The move from static benchmark scores to agent and multimodal testing reflects a broader industry shift from model capability to task execution. - By focusing on Taiwan-specific context, iKala is drawing a line between language coverage and real-world usability.

What’s next: - iKala wants TMMLU+ to serve as a shared foundation for defining AI capability and building trust across Taiwan’s AI ecosystem. - The company expects future evaluation to center more on local data, workflow fit and measurable business outcomes. - As AI systems evolve, iKala says Taiwan’s advantage will depend on its ability to independently measure and validate those systems.

Disclaimer: This article was produced by AGP Wire with the assistance of artificial intelligence based on original source content and has been refined to improve clarity, structure, and readability. This content is provided on an “as is” basis. While care has been taken in its preparation, it may contain inaccuracies or omissions, and readers should consult the original source and independently verify key information where appropriate. This content is for informational purposes only and does not constitute legal, financial, investment, or other professional advice.

Sign up for:

The Asia Reporter

The daily local news briefing you can trust. Every day. Subscribe now.

By signing up, you agree to our Terms & Conditions.

Share this page:

Advanced Search Options

Search for:

Search scope:

Type:

Search in:

Date range:

The last

Sort by:

Sign up for:

The Asia Reporter

The daily local news briefing you can trust. Every day. Subscribe now.

By signing up, you agree to our Terms & Conditions.