Most AI coverage aimed at hedge funds asks which vendor to buy. But for a fund with its own quants and its own data, the correct question is rather narrower and harder: if your engineers are building proprietary alpha tooling themselves, which frontier model provider should sit underneath it?
Neither Anthropic nor OpenAI appears in the Qaike vendor directory, and that is deliberate. They are not fund-tech vendors. They sit one layer down, underneath the vendors you buy and underneath the tools you build. This piece looks only at the second case.
On 11 February 2026, Man Group announced a partnership with Anthropic spanning its investment process, distribution and HR functions. Man runs roughly $214bn.
Man was not buying a finance chatbot. It already had AlphaGPT, its own proprietary alpha tooling. CTO Gary Collier framed LLMs as having accelerated our ability to create proprietary alpha generation tools. The products named in the announcement were Claude Skills and Claude Code, which are the building blocks for in-house technology.
The same pattern shows in Anthropic other buy-side references. Walleye Capital cited more capable engineering in fewer turns. IMC cited trading-analysis evaluations. Citadel cited Excel.
OpenAI has no publicly announced equivalent. Its closest buy-side association is Bridgewater AIA Labs, a roughly $2bn effort that reportedly draws on OpenAI, Anthropic and Perplexity together. That is a multi-provider architecture, not an endorsement.
Building alpha tooling in-house is a different workload from analyst assistance via a chatbot. Five requirements dominate.
Scores below are BenchLM category composites as at 27 September 2026, not single benchmarks. Each category pools dozens of underlying evals onto a calibrated scale, and the methodology is published. Follow the links to see what sits inside each one.
| Measure | Claude Opus 5.5 | Claude Opus 5 | GPT-6 Astra |
|---|---|---|---|
| Agentic | 87.9 | 77.3 | 70.6 |
| Coding | 87.6 | 72.8 | 74.6 |
| Reasoning | 82.4 | 77.3 | 89.5 |
| Finance Agent v2 | 58.6% | 58.6% | 53.5% |
| SWE-bench Pro | 89.9% | 79.2% | not listed |
| Input price, $/1M | 4 | 5 | 10 |
| Output price, $/1M | 20 | 25 | 50 |
| Context window | 1M | 1M | 1.05M |
| Requirement | Anthropic | OpenAI |
|---|---|---|
| Cheapest usable tier | Sonnet 5 at $2 in, $10 out | GPT-5.6 Luna at $0.20 in, $1.20 out |
| Coding agent | Claude Code | Codex |
| Clouds | AWS, Google Cloud, Azure via Foundry | Azure, OpenAI direct |
| UK data residency | Not on first-party. Via Bedrock-EU or Vertex-EU | Yes, UK named explicitly |
| Certifications | SOC 2 Type II, ISO 27001, ISO 42001 | SOC 2, FedRAMP via Azure Government |
| Named buy-side build references | Man Group, Walleye, IMC, Citadel, Balyasny | None announced |
Agentic work is the clearest gap, and it widened. Opus 5.5 scores 87.9 against Astra 70.6, a 17.3 point lead where Opus 5 held 6.7. That matters more than it looks, because alpha tooling is mostly long-running agent loops rather than single answers. Reliability compounds on a run that pulls a dataset, writes code, executes it, checks the risk model and revises is exactly. The category is weighted towards computer use, terminal work and browser research, which is a fair proxy for that shape of job.
Coding is the reversal worth noticing. Opus 5 has slipped behind Astra, 72.8 to 74.6. Opus 5.5 then jumps to 87.6, and takes SWE-bench Pro at 89.9%. Astra does not appear on that leaderboard at all. So the coding case for Anthropic now rests entirely on the newest model rather than the line-up as a whole.
Price moved the right way too. Opus 5.5 runs at $4 in and $20 out, undercutting Opus 5 and coming in at 60% below Astra, while scoring higher on everything except reasoning. That is an unusual direction of travel and it is the single strongest argument in Anthropic favour this quarter.
Three places, and each is decisive in the right circumstances.
Reasoning. Astra 89.5 against Opus 5.5 82.4 is the one category OpenAI still holds, and it holds it clearly. The category leans on long-context retrieval and abstract reasoning. If your tooling turns on hard quantitative work rather than long tool-use chains, that is the more relevant score.
Volume economics. GPT-5.6 Luna scores 55.0% on Finance Agent v2 at $0.20 per million input tokens. It costs a fraction of anything comparable at Anthropic, whose cheapest useful tier is ten times the input price. For screening tens of thousands of filings, transcripts or securities, there is no Anthropic equivalent.
The finance benchmark has not moved. This is the caveat that should temper the rest. Opus 5.5 scores 58.6% on Finance Agent v2, identical to Opus 5 and slightly below Fable 5.1. A model can gain ten points of general agentic capability and none of it shows up in financial analyst tasks. Take the headline numbers as a signal about engineering throughput, not about analysis quality.
One finding cuts across both. Anthropic Model Context Protocol has become the de-facto integration standard, and the vendor ecosystem has settled on it. Over August and September 2026 alone, Arcesium published an MCP for private markets, Blueflame AI shipped an MCP connections library for proprietary systems, and Deal Engine argued that a single connection to Claude pays for itself.
For a fund building in-house, this is the most useful fact in the piece. Connector work done against MCP largely transfers between providers. The cost of switching model providers is now much lower than the cost of switching your data plumbing. Build the plumbing once, deliberately, and treat the model as replaceable.
Anthropic is the stronger default for proprietary alpha tooling today, and Opus 5.5 strengthened that position rather than merely maintaining it. It leads on agentic reliability by a wide margin, leads on coding, costs less than the model it replaced, and Anthropic is still the only one of the two with named hedge fund build references.
Three exceptions are worth taking seriously. Route high-volume, low-complexity work to OpenAI cheap tier, because the price gap is too wide to ignore. If UK residency on a first-party platform is a hard compliance requirement rather than a preference, OpenAI solves it out of the box. And if your workload is reasoning-heavy rather than agentic, Astra still wins that category outright.
There is also a lesson about how fast each relative capabilities evolve in this space. As Auquan put it in August, the expensive mistake is building permanently around capabilities that will not stay fixed.
To see how the vendors building on top of these models compare for specific fund workflows, browse the full Qaike vendor directory.
This article is based on publicly available information and reflects Qaike own analysis and opinions. It is not intended to provide professional advice, endorsement, or a definitive assessment of any vendor or product. While every effort has been made to ensure accuracy, completeness and correctness cannot be guaranteed. Benchmark figures are third-party composites and were current on 27 September 2026.
Created by Qaike – Powered by AI.