TC Labs at FinMMEval 2026: Putting Specialized Multi-Agent AI to the Test in Finance



Can AI agents think like a team of financial analysts — and hold up when the questions get hard, multilingual, and market-moving? That's the question TC Labs, Trading Central's research unit, set out to test at the FinMMEval Lab 2026, part of the CLEF 2026 conference in Jena, Germany.
Our data scientists, Elvys Linhares Pontes and Mohamed Benjannet, entered all three of the lab's tasks: (1) Financial Exam Question Answering, (2) Multilingual Financial Question Answering, and (3) Financial Decision Making. Their paper, TCLabs @ FinMMEval 2026: Systems for Tasks 1, 2 and 3, is now published in the CEUR Workshop Proceedings.
The competition and our findings
Across all three tasks, TC Labs took the same approach: instead of asking a single large language model to do everything, it split each problem across specialized AI agents, much like a research desk divides work among analysts.
- Task 1, Financial Exam Q&A: Multiple-choice questions at professional-certification level, spanning accounting, valuation, ethics, regulation, and portfolio management, in Arabic, Chinese, English, and Hindi. TC Labs' panel of specialist agents, checked by a critic and a judge, scored 90% in Hindi (3rd place) and 89.5% in Arabic (5th place).
- Task 2, Multilingual Financial Q&A: Concise, evidence-backed answers drawn from 10-K and 10-Q filings plus news in several languages. TC Labs' separate agents for news and financial statements cross-checked each other's findings, placing 9th, with answer precision as the clear next step.
- Task 3, Financial Decision Making: A daily BUY, HOLD, or SELL call on Tesla (TSLA) and Bitcoin (BTC), with a three-minute limit per decision. For TC Labs, adding Trading Central news and technical indicators improved every model tested. In a falling market, the live agent outperformed buy-and-hold by 13.55 percentage points on Tesla and 15.57 points on Bitcoin (leaderboard as of July 3, 2026).

What it means for brokers and financial institutions
As a broker, you see the same challenges on your platform every day: clients asking complex questions, in many languages, about markets that move on the news. Competitions like this one show us what actually works, so we can keep building solutions that match what your clients need.
Quality data beats bigger models
In Task 3, adding Trading Central news and technical indicators improved every model tested, while extra training did not. AI is only as good as the research underneath it, and that's where expert-backed analytics earn their keep.
Trading Central’s MCP Server enables brokers to ship AI agents fueled by institutional-grade market data and analytics in a single integration. See it in action!
Specialization builds trust
Splitting a question across expert agents, with a built-in critic, makes the reasoning easier to check and harder to derail. That matters when your clients act on what they read.
Our one-stop shop of market research offers specialized tools across all analytical domains: technicals, fundamentals, macro, news & sentiment, volatility, and more. Get a demo to see how these tools can drive usage and trades on your platform!
Multilingual is non-negotiable
Strong results in Arabic and Hindi show that sophisticated financial reasoning doesn't have to be English-first, a real advantage for brokers serving global client bases.
Trading Central research is available in up to 32 languages, delivered by a global team of expert market analysts in Paris, Hong Kong, and Ottawa to empower brokers around the globe. Let’s discuss how our team can help yours!
Read the full paper
TC Labs will keep pushing on what makes financial AI accurate, transparent, and reliable to investors. For the full methodology, prompts, and leaderboards, read the paper in the CEUR Workshop Proceedings.
About FinMMEval Lab at CLEF 2026
CLEF (the Conference and Labs of the Evaluation Forum) is a long-running European forum where research teams test AI systems head-to-head on shared tasks, with common datasets and public leaderboards. The 2026 edition ran September 21–24 in Jena, Germany.
FinMMEval introduces the first multilingual and multimodal evaluation framework for financial large language models, assessing models’ abilities in understanding, reasoning, and decision-making across languages and modalities to promote robust, transparent, and globally inclusive financial AI systems.