Between June and October 2026, vendors serving hedge funds and private markets firms published a steady run of benchmarks. Leaderboards, open datasets, head-to-head tests and even cash guarantees all claim to show which tool you can trust with real work.
That is useful but also a problem. Every test below was built by a vendor, and in two cases that vendor’s own system comes out on top. This digest looks at what each one measures, what it claims, and how much weight it can bear. The short answer: several score a vendor’s data, tools and at times its own models, not a bare model on equal terms.
Most finance AI vendors do not build their own large language model (LLM). They build a harness: the retrieval, data feeds, tools, prompts and checks that wrap around frontier models from Anthropic, OpenAI, Google and others. So when a chart shows “Vendor X” beating “Claude” or “GPT”, it usually compares a full product built on a frontier model against the same kind of model with fewer tools. The benchmark then mostly shows whether the harness adds value.
The benchmarks below test three different things:
| What is scored | Meaning | Benchmarks here |
|---|---|---|
| Bare or lightly equipped models | Frontier and open models, same tools for all | Aiera Leaderboard, Rogo BFB (set-up not disclosed) |
| A model plus a vendor’s data | Same models, with and without the vendor’s data feed | Aiera Lift |
| A vendor product vs generic set-ups | Vendor’s full system against frontier models with basic tools | FrontierFinance (Samaya entry), Imprima |
Below is what each vendor’s product runs on, as far as they disclose:
| Vendor | Under the hood |
|---|---|
| Aiera | Data and an MCP server, not an AI model. Frontier models do the reasoning |
| Samaya | Its own “custom models, data index and retrieval engines”. How these relate to frontier models is not disclosed |
| Rogo | Frontier models from Anthropic and others, chosen per task |
| Imprima | An LLM customised for redaction (per a 2023 post). The base model is not named |
| Hudson Labs | A mix of OpenAI and five or more other models, plus its own retrieval and checks |
| Blueflame AI | Model-agnostic harness over Anthropic, OpenAI, Google and other models |
| Marvin Labs | Anthropic and OpenAI models, per its security page |
| Benchmark | Publisher | Published | What it tests | Open? | Headline result |
|---|---|---|---|---|---|
| Aiera Leaderboard | Aiera | 17 Jun 2026 | 16 models on research, Q&A, summaries, sentiment | Partly | Best research score 46.9 out of 100 |
| Aiera Lift | Aiera | 24 Jun 2026 | Gain from adding Aiera’s data via MCP | No | Average score up from 13% to 32% |
| FrontierFinance | Samaya | 9 Jul 2026 | 220 expert-written investor tasks | Yes | Top score 50.8%, Samaya’s own system |
| Big Finance Bench (BFB) | Rogo | 22 Sep 2026 | Valuation, statement analysis, forecasting | No | No scores disclosed |
| Smart Redaction test | Imprima | 1 Oct 2026 | Redaction of names, addresses, firms in 80 documents | No | F1 of 0.84 vs 0.78 (GPT-5.6) and 0.66 (Azure) |
The benchmarks differ in what they ask, how answers are marked and what the final number means. In brief:
| Benchmark | Tasks | How answers are marked | What the score means |
|---|---|---|---|
| Aiera Leaderboard | About 150 analyst-style research questions, answered using Aiera’s data, plus Q&A, summary and sentiment tasks | A model from a different family checks each answer fact by fact against the expected facts | 0 to 100 per task, with research weighted at 60% of the overall score |
| Aiera Lift | 150 research questions, each run twice: with and without Aiera’s data | A separate judge model checks each answer against the expected facts | Average share of expected facts found, before and after adding the data |
| FrontierFinance | 220 open-ended tasks across six workflows, from screening to modelling, written by finance experts | Each task has a rubric (11,543 points in total, some marked must-have). Three judge models vote on each point | Share of rubric points an answer meets, averaged across all tasks |
| Big Finance Bench | Valuation, statement analysis, forecasting and other workflows | Finance practitioners mark correctness, choice of sources, definitions and execution | Not published |
| Imprima redaction test | 80 due diligence documents in seven languages | Each tool’s redactions are compared with a reference set of sensitive items | Recall, precision and F1 (see below) |
A few terms used in these results:
Aiera launched its Leaderboard in June. It scores 16 third-party models, closed and open, across four tasks, with research questions carrying 60% of the weight. For research, every model connects to the same Aiera data server and nothing else. That makes it the closest thing here to a fair model-against-model test. A grader from a different model family checks each answer fact by fact.
The results are sobering. The leader, Claude Opus 4.6, scored 46.9 on research. Only five models cleared 30, and seven scored below 10. Aiera also found that research skill tracks general financial knowledge only weakly. A model that does well on finance trivia may still fail at analyst work.
A week later came the Aiera Lift. It runs 13 models twice: once with their own knowledge and web search, once with Aiera’s data. The average score rose from 13% to 32%, but only six models improved in a meaningful way. Aiera’s own phrase is that not every model “can convert access into better responses”.
Three caveats apply. The research questions stay private for licensing reasons, so no one can reproduce the results. The Lift study measures the value of the product Aiera sells. And in Lift, the set-up was not uniform: frontier models used their providers’ own search and connectors, while open models shared a search tool Aiera admits was weaker.
Samaya released FrontierFinance in July wirh 220 open-ended tasks across six workflows, from screening to modelling, each marked against a detailed rubric. The dataset, rubrics and code are public, which makes it the only fully open benchmark here.
The best scores sit around half marks. Samaya’s in-house system leads at 50.8%, ahead of Claude Fable 5 at 49.2% and Claude Opus 4.8 at 45.0% while being roughly 4 times cheaper.
The caveat is that Samaya’s entry is its full product. The frontier models ran with either their built-in web search or a common six-tool agent set-up.
More precisely three distinct types of harness were considered
Every harness connects only to publicly available data, such as web content, public filings, and news, and none draw on private sources such as proprietary research.
Rogo announced in September that its Big Finance Bench now feeds an index run by Fireworks AI. Finance practitioners built and grade it, covering valuation, statement analysis and forecasting. Rogo itself runs on frontier models, and says it uses BFB to route each task to the model with the best mix of quality and cost. So BFB scores models for Rogo’s harness, not Rogo against them. The announcement gives no task count, no model list and no scores.
Imprima tested its Smart Redaction tool against Microsoft Azure Redaction (used by most data rooms) and GPT-5.6 on 80 documents in seven languages. It reports an F1 score of 0.84, against 0.78 for GPT-5.6 and 0.66 for Azure, and claims to run about five times faster than GPT-5.6.
Again, note what is compared. Smart Redaction appears to be a model Imprima has tuned for this one task, judging by a 2023 post. GPT-5.6 ran bare, with a prompt and no task-specific training. If so, a tuned specialist beating a general model on a narrow job is expected. The more useful signal is how close the general model came.
Imprima’s August post is more telling. It made its own test harder, scoring how much of each sensitive item was caught rather than whether any part was. That is the right instinct for a task where one missed name can breach GDPR. The tests remain in-house, though, with no public dataset.
Hudson Labs took a different route. Its No Hallucination Guarantee pays $50 for any figure its Co-Analyst gives that cannot be traced to a filing or call. Co-Analyst blends OpenAI and several other models with Hudson Labs’ own retrieval and checks, so the guarantee covers the whole system, not one model. The exclusions are wide: calculations, omissions, interpretation and uploaded documents all fall outside it, and Hudson Labs judges each claim itself. At $50 a claim the sum is small, and the scope is narrow. Still, it is the only vendor here putting money behind traceability, the quality firms most need.
Two vendors argue that scores miss the point. Blueflame AI, whose platform sits over models from several providers, contends that funds should test whole systems, including retrieval, tools and refusals, not models alone. It reports seeing the same model perform very differently in different harnesses. Marvin Labs, which also runs on frontier models, argues that models now answer well once they find the right material. The real test, it says, is which sources a tool trusts and how much weight it gives each.
Both points are fair, but both firms sell a harness over other providers’ models, so the view that the system matters more than the model also suits their business.
Four points stand out.
Vendor benchmarks are a starting point, not proof. Before you buy, test the tool on your own work.
| Question | Why it matters |
|---|---|
| Is the score for a model or a full system? | A product beating a bare model proves the harness, not the model |
| Which models sit under the product, and can we switch them? | You may already pay for the same model elsewhere |
| Who built the benchmark, and does their product win it? | All tests here are vendor-made |
| Can we see the questions and grading? | Only one benchmark is fully open |
| Can every figure be traced to a source? | Traceability is what auditors and ICs need |
| How does it score on 30 to 50 of our own past tasks, against our own Claude or ChatGPT set-up? | Your workflows, and your fallback option, are the only fair comparison. Thirty to 50 tasks is enough to show a pattern, yet small enough for your team to mark by hand |
The vendors above are all listed in Qaike’s vendor directory, where you can compare them by function and use case.
This article is based on publicly available information and reflects Qaike’s own analysis and opinions. It is not intended to provide professional advice, endorsement, or a definitive assessment of any vendor or product. While every effort has been made to ensure accuracy, completeness and correctness cannot be guaranteed.
Created by Qaike – Powered by AI.