Skip to main content

Qaike

Finance AI Benchmarks: What They Really Test

Rendered race track with parallel lanes of blocks and robots heading to a finish gantry, illustrating finance AI tools being benchmarked against each other

Between June and October 2026, vendors serving hedge funds and private markets firms published a steady run of benchmarks. Leaderboards, open datasets, head-to-head tests and even cash guarantees all claim to show which tool you can trust with real work.

That is useful but also a problem. Every test below was built by a vendor, and in two cases that vendor’s own system comes out on top. This digest looks at what each one measures, what it claims, and how much weight it can bear. The short answer: several score a vendor’s data, tools and at times its own models, not a bare model on equal terms.

What the benchmarks really test

Most finance AI vendors do not build their own large language model (LLM). They build a harness: the retrieval, data feeds, tools, prompts and checks that wrap around frontier models from Anthropic, OpenAI, Google and others. So when a chart shows “Vendor X” beating “Claude” or “GPT”, it usually compares a full product built on a frontier model against the same kind of model with fewer tools. The benchmark then mostly shows whether the harness adds value.

The benchmarks below test three different things:

What is scoredMeaningBenchmarks here
Bare or lightly equipped modelsFrontier and open models, same tools for allAiera Leaderboard, Rogo BFB (set-up not disclosed)
A model plus a vendor’s dataSame models, with and without the vendor’s data feedAiera Lift
A vendor product vs generic set-upsVendor’s full system against frontier models with basic toolsFrontierFinance (Samaya entry), Imprima

Below is what each vendor’s product runs on, as far as they disclose:

VendorUnder the hood
AieraData and an MCP server, not an AI model. Frontier models do the reasoning
SamayaIts own “custom models, data index and retrieval engines”. How these relate to frontier models is not disclosed
RogoFrontier models from Anthropic and others, chosen per task
ImprimaAn LLM customised for redaction (per a 2023 post). The base model is not named
Hudson LabsA mix of OpenAI and five or more other models, plus its own retrieval and checks
Blueflame AIModel-agnostic harness over Anthropic, OpenAI, Google and other models
Marvin LabsAnthropic and OpenAI models, per its security page

The benchmarks at a glance

BenchmarkPublisherPublishedWhat it testsOpen?Headline result
Aiera LeaderboardAiera17 Jun 202616 models on research, Q&A, summaries, sentimentPartlyBest research score 46.9 out of 100
Aiera LiftAiera24 Jun 2026Gain from adding Aiera’s data via MCPNoAverage score up from 13% to 32%
FrontierFinanceSamaya9 Jul 2026220 expert-written investor tasksYesTop score 50.8%, Samaya’s own system
Big Finance Bench (BFB)Rogo22 Sep 2026Valuation, statement analysis, forecastingNoNo scores disclosed
Smart Redaction testImprima1 Oct 2026Redaction of names, addresses, firms in 80 documentsNoF1 of 0.84 vs 0.78 (GPT-5.6) and 0.66 (Azure)

How the tests work

The benchmarks differ in what they ask, how answers are marked and what the final number means. In brief:

BenchmarkTasksHow answers are markedWhat the score means
Aiera LeaderboardAbout 150 analyst-style research questions, answered using Aiera’s data, plus Q&A, summary and sentiment tasksA model from a different family checks each answer fact by fact against the expected facts0 to 100 per task, with research weighted at 60% of the overall score
Aiera Lift150 research questions, each run twice: with and without Aiera’s dataA separate judge model checks each answer against the expected factsAverage share of expected facts found, before and after adding the data
FrontierFinance220 open-ended tasks across six workflows, from screening to modelling, written by finance expertsEach task has a rubric (11,543 points in total, some marked must-have). Three judge models vote on each pointShare of rubric points an answer meets, averaged across all tasks
Big Finance BenchValuation, statement analysis, forecasting and other workflowsFinance practitioners mark correctness, choice of sources, definitions and executionNot published
Imprima redaction test80 due diligence documents in seven languagesEach tool’s redactions are compared with a reference set of sensitive itemsRecall, precision and F1 (see below)

A few terms used in these results:

  • Judge model. An AI model used to mark answers. It is faster and cheaper than human marking, but it can share blind spots with the model it marks. That is why Aiera uses a judge from a different model family and Samaya uses three.
  • Rubric. A checklist of points a good answer should contain. The score is the share of points met, so a 50% score means half the expected content, not half the answers right.
  • Recall. The share of items that should have been found that the tool actually found. In redaction, every missed name lowers recall.
  • Precision. The share of what the tool flagged that was correct. Low precision means too much is blacked out.
  • F1. A single score that balances recall and precision, from 0 to 1. It hides which of the two is weaker, so the underlying figures matter.

Aiera: like for like, then a test of its own data

Aiera launched its Leaderboard in June. It scores 16 third-party models, closed and open, across four tasks, with research questions carrying 60% of the weight. For research, every model connects to the same Aiera data server and nothing else. That makes it the closest thing here to a fair model-against-model test. A grader from a different model family checks each answer fact by fact.

The results are sobering. The leader, Claude Opus 4.6, scored 46.9 on research. Only five models cleared 30, and seven scored below 10. Aiera also found that research skill tracks general financial knowledge only weakly. A model that does well on finance trivia may still fail at analyst work.

A week later came the Aiera Lift. It runs 13 models twice: once with their own knowledge and web search, once with Aiera’s data. The average score rose from 13% to 32%, but only six models improved in a meaningful way. Aiera’s own phrase is that not every model “can convert access into better responses”.

Three caveats apply. The research questions stay private for licensing reasons, so no one can reproduce the results. The Lift study measures the value of the product Aiera sells. And in Lift, the set-up was not uniform: frontier models used their providers’ own search and connectors, while open models shared a search tool Aiera admits was weaker.

Samaya: the most open dataset, but a product among models

Samaya released FrontierFinance in July wirh 220 open-ended tasks across six workflows, from screening to modelling, each marked against a detailed rubric. The dataset, rubrics and code are public, which makes it the only fully open benchmark here.

The best scores sit around half marks. Samaya’s in-house system leads at 50.8%, ahead of Claude Fable 5 at 49.2% and Claude Opus 4.8 at 45.0% while being roughly 4 times cheaper.

The caveat is that Samaya’s entry is its full product. The frontier models ran with either their built-in web search or a common six-tool agent set-up.

More precisely three distinct types of harness were considered

  1. Web search harness: an agentic frontier model paired with their built-in web search API.
  2. The open-source Finance Agent v2 harness: an agentic model connected to six specialized tools built for finance tasks, covering the SEC EDGAR API, a market price data API, web search, HTML parsing, search within long HTML content, and a calculator.
  3. The in-house Samaya agent harness: a more sophisticated harness that combines Samaya’s custom models, data index, and retrieval engines, optimized for both quality and efficiency.

 

Every harness connects only to publicly available data, such as web content, public filings, and news, and none draw on private sources such as proprietary research.

Source: https://samaya.ai/blog/frontier-finance

Rogo: a benchmark for picking models

Rogo announced in September that its Big Finance Bench now feeds an index run by Fireworks AI. Finance practitioners built and grade it, covering valuation, statement analysis and forecasting. Rogo itself runs on frontier models, and says it uses BFB to route each task to the model with the best mix of quality and cost. So BFB scores models for Rogo’s harness, not Rogo against them. The announcement gives no task count, no model list and no scores.

Imprima: a narrow task, measured with care

Imprima tested its Smart Redaction tool against Microsoft Azure Redaction (used by most data rooms) and GPT-5.6 on 80 documents in seven languages. It reports an F1 score of 0.84, against 0.78 for GPT-5.6 and 0.66 for Azure, and claims to run about five times faster than GPT-5.6.

Again, note what is compared. Smart Redaction appears to be a model Imprima has tuned for this one task, judging by a 2023 post. GPT-5.6 ran bare, with a prompt and no task-specific training. If so, a tuned specialist beating a general model on a narrow job is expected. The more useful signal is how close the general model came.

Imprima’s August post is more telling. It made its own test harder, scoring how much of each sensitive item was caught rather than whether any part was. That is the right instinct for a task where one missed name can breach GDPR. The tests remain in-house, though, with no public dataset.

Hudson Labs: a guarantee instead of a score

Hudson Labs took a different route. Its No Hallucination Guarantee pays $50 for any figure its Co-Analyst gives that cannot be traced to a filing or call. Co-Analyst blends OpenAI and several other models with Hudson Labs’ own retrieval and checks, so the guarantee covers the whole system, not one model. The exclusions are wide: calculations, omissions, interpretation and uploaded documents all fall outside it, and Hudson Labs judges each claim itself. At $50 a claim the sum is small, and the scope is narrow. Still, it is the only vendor here putting money behind traceability, the quality firms most need.

What other vendors say

Two vendors argue that scores miss the point. Blueflame AI, whose platform sits over models from several providers, contends that funds should test whole systems, including retrieval, tools and refusals, not models alone. It reports seeing the same model perform very differently in different harnesses. Marvin Labs, which also runs on frontier models, argues that models now answer well once they find the right material. The real test, it says, is which sources a tool trusts and how much weight it gives each.

Both points are fair, but both firms sell a harness over other providers’ models, so the view that the system matters more than the model also suits their business.

What this tells us

Four points stand out.

  • Top scores are about half marks. The best research score on Aiera’s Leaderboard is 46.9 out of 100, and the best system on FrontierFinance meets 50.8% of rubric points. The two scales measure different things, but both say the same: no tool is close to doing analyst work unchecked.
  • Most “vendor vs model” charts compare a harness to a model. Since most vendors run on frontier models, a win mostly shows the value of their data, tools and checks, not a better model. That value is real, but it is a different claim. The two vendors that top their own tests, Samaya and Imprima, are the exceptions: both use their own or tuned models, and neither discloses enough to show how much of the win comes from the model and how much from the rest.
  • A rank on one test says little about another. Each test uses its own tasks, tools, data access and scale. Even within one test, Aiera found that a model’s general finance knowledge predicts its research skill only weakly. Compare tools within one test, not across tests.
  • Good sources are necessary, but not sufficient. Aiera, Hudson Labs and Marvin Labs all stress what a tool retrieves and whether it can show where a number came from. Yet Aiera’s own data cuts the other way too: with the same data, research scores ranged from 46.9 to below 10, and only six of 13 models gained much from Aiera’s feed. This shows that the model must still make good use of quality sources, which can vary across models.

How to judge a tool for your firm

Vendor benchmarks are a starting point, not proof. Before you buy, test the tool on your own work.

QuestionWhy it matters
Is the score for a model or a full system?A product beating a bare model proves the harness, not the model
Which models sit under the product, and can we switch them?You may already pay for the same model elsewhere
Who built the benchmark, and does their product win it?All tests here are vendor-made
Can we see the questions and grading?Only one benchmark is fully open
Can every figure be traced to a source?Traceability is what auditors and ICs need
How does it score on 30 to 50 of our own past tasks, against our own Claude or ChatGPT set-up?Your workflows, and your fallback option, are the only fair comparison. Thirty to 50 tasks is enough to show a pattern, yet small enough for your team to mark by hand

The vendors above are all listed in Qaike’s vendor directory, where you can compare them by function and use case.

This article is based on publicly available information and reflects Qaike’s own analysis and opinions. It is not intended to provide professional advice, endorsement, or a definitive assessment of any vendor or product. While every effort has been made to ensure accuracy, completeness and correctness cannot be guaranteed.

Created by Qaike – Powered by AI.