Most fund teams now use or plan to AI somewhere. Some have live workflows that sort investor emails, read CIMs or check documents. Many more are testing, or planning their first. In almost every case the engine is one or more large language models (LLMs), the kind of model behind ChatGPT, Claude or Gemini.
LLMs are very good at reading and writing. But much of a fund’s operations work isn’t about generating text. It is a stream of small judgement calls: which team, which cause, approve or escalate, in scope or not. Using an LLM for each one is slow, costly and hard to explain to an auditor.
A new kind of model, the decision model, does these calls and nothing else. This guide explains what it is, where it fits and how to govern it, whether you are reviewing live workflows or planning your first.
For a fund starting out, the natural default is to point an LLM at every task. It works well in a demo. In production, a few issues tend to appear.
| Issue | What it looks like in operations |
|---|---|
| Invented answers | The model states a wrong fact with full confidence. The FINOS AI Governance Framework notes there is still no reliable way to remove this. |
| Different answers to the same case | Run the same case twice and get two outcomes. FINOS lists this as a separate risk. |
| Cost that grows with volume | You pay for every word in and out, so checking every item, not a sample, gets expensive fast. |
| Speed | The model writes its answer word by word. That takes seconds, and agent tasks can take many minutes. Too slow to sit inside a live process. |
| Weak audit trail | The decision is buried in prose. It is hard to test, replay or show to a regulator. |
| No sense of doubt | The model rarely says how sure it is, so you can’t tell safe cases from risky ones. |
None of this makes LLMs the wrong tool. It makes them the wrong tool for every step.
A decision model reads the facts of a case and a question, then picks from options you set. It doesn’t write. It returns each option with a probability or confidence score for the answer.
Compare asking a colleague to write a memo on why a trade broke with handing them a form with four boxes: timing, price, quantity, missing booking. The LLM writes the memo. The decision model ticks a box and tells you how sure it is.
| LLM | Decision model | |
|---|---|---|
| Job | Read and write | Judge and choose |
| Output | Free text | A choice from your list, with a confidence score |
| Can it invent an answer? | Yes | No: it can only pick from your options |
| Speed | Seconds to minutes | 70 to 500 milliseconds (TypeSafe) |
| Cost | Charged for input and output | Charged for input only, at a small fraction of LLM rates |
| Audit | Prose you have to interpret | Structured data you can log and replay |
| Knows when it is unsure? | Rarely | Yes, by design: higher confidence should mean higher accuracy |
One point needs care. Decision models are not fully deterministic either: the same case can come back with slightly different scores (Firecrawl). The predictability comes from the design around the model. The answer can only be one of your options, and fixed rules in your own systems decide what happens next. Same score, same action, every time.
Efficiency. When a check costs a fraction of a cent and takes under a second, you can check every item rather than a sample. Early users report large gains. Vercel saw results 5 to 18 times faster, and more accurate. Bryo AI found costs 10 to 20 times lower than with Gemini (TechCrunch). LLM spend then goes only to work that needs reading and writing.
Predictability. The model can only answer from options you wrote. Thresholds that the business signs off decide what happens next. That makes a process easier to test before go-live and easier to explain afterwards.
Governance. Each call leaves a record: the facts, the question, the options, the scores and the action taken. It is data, not prose, so you can replay it, sample it and report on it. When an operational due diligence questionnaire asks “how does the AI make decisions, and who checks it?”, you can answer with evidence.
The pattern is the same each time. The decision model sits at the point where a workflow splits, and sends each case down one of three paths:
This works whether you are adding it to a live workflow or designing one from scratch.
Routine requests, such as a capital call date or a copy of a report, go straight to the right team with a template. Requests that need a tailored reply go to an LLM for a draft, which investor relations approves. Complaints, and anything the model is unsure about, go to a senior person. The LLM now handles only the emails that need writing.
The decision model checks each CIM or teaser against the fund’s criteria: sector, region, size and red flags. Clear misses are parked on a list that the deal team checks each week, so nothing is dropped unseen. Clear fits go to an LLM to draft a first screening memo. Borderline cases go to the deal team.
Each morning, analysts sort breaks by hand. A decision model can sort them first by likely cause (timing, price, quantity, missing booking) and flag known patterns. Known timing breaks are rechecked the next day and closed if matched. Breaks that need work go to an LLM, which gathers the trade data and drafts a note. Material or unclear breaks go straight to an analyst.
Note what stays out of the model. Amounts, dates and tolerances are checked by the reconciliation system’s own rules, because decision models are weak at maths and dates (Firecrawl).
As funds move from chat assistants to agents that send emails, update records or move files, each planned action needs a check. A decision model can ask, in under a second: is this hard to undo, does it go beyond the task, does it touch client data? One open-source test of this approach held 42 risky actions across about 17,000 calls (Firecrawl). Hard rules, such as allow-lists and spending limits, still apply whatever the model says.
| If the step is… | Use |
|---|---|
| Picking from a known list (route, sort, approve, flag) | Decision model |
| Scoring against a scale (urgency, risk, fit) | Decision model |
| Writing, summarising or explaining | LLM |
| Research across many sources, or multi-step tasks | LLM or agent |
| Maths, dates, limits, tolerances | Plain code |
| High impact or hard to undo | A person, with the model’s score as input |
Every decision comes with a confidence score, so you choose when the system acts alone. Set that line by the cost of being wrong, not by habit. The thresholds below are examples only.
| Risk of error | Example | Act automatically when confidence is at least |
|---|---|---|
| Low | Routing an investor email | 0.80 |
| Medium | Parking a deal that misses the criteria | 0.90 |
| High | Releasing a payment or closing a material break | Never: always a person |
Tune these from pilot data. Avoid setting a cut-off where many scores cluster, since small shifts in score would then flip outcomes. Review the thresholds each quarter.
| Control | Why it matters |
|---|---|
| Fix the model version | A silent update to the vendor’s latest model changes behaviour (VentureBeat) |
| Log every call in full | Facts, question, options, scores and action taken, so any decision can be replayed |
| Keep hard rules in code | Limits and allow-lists must not depend on a model’s judgement |
| Keep a person on high-impact steps | Whatever the confidence score says |
| Test against planted text | In one test, a fake “approved” field cut the score for blocking a harmful command from 0.76 to 0.48 (VentureBeat) |
| Test with options reordered | If the answer changes, the question needs rewording |
| Name an owner for each threshold | Someone in the business signs off, not only IT |
| Review each quarter | Check error rates by path and adjust thresholds |
Standards such as SOC 2 and ISO 27001 do not yet set out how to audit decisions like these (VentureBeat). Writing your own approach down now puts you ahead of the question when investors or regulators ask it.
The category took off with Jev, launched in early access on 15 September 2026 by TypeSafe AI. Its founder, Diogo Almeida, is a former OpenAI researcher who worked on ChatGPT (TechCrunch). Jev costs $0.042 per million input tokens, and output is free. TypeSafe claims it is 193 times faster and 445 times cheaper than OpenAI’s GPT-6 Astra (Tom’s Hardware). For scale, Astra is priced at $10 per million input tokens and $50 per million output tokens (DataCamp). Direct sign-ups were paused on 22 September because of demand, but Jev is also offered through Vercel, OpenRouter and Cloudflare (Firecrawl, eesel).
Open-weight decision models, which a fund can run on its own servers, followed within days. The most complete is Laya from ConvAI Innovations: free under an Apache 2.0 licence, covering more than 100 languages, and taking about 7 to 33 milliseconds per question on the firm’s own hardware (Medium). Others, such as Kev, Von and NanoJev, are early projects best kept to trials (DataCamp).
| Option | What it is | Where it runs | Pricing | Confidence scores | Best fit |
|---|---|---|---|---|---|
| Jev (TypeSafe AI) | Decision model | Vendor-hosted | $0.042 per million input tokens, output free | Calibrated, built in | High-volume decisions, fast start |
| Laya (ConvAI Innovations) | Decision model, open weights | Your own servers | Free licence, you pay for hosting | Built in | Data must stay in-house, several languages |
| Kev, Von, NanoJev | Open-source decision models | Your own servers | Free licence, you pay for hosting | Varies | Trials and research |
| LLM with structured output (OpenAI, Gemini, Claude) | LLM forced to answer in a set format | Vendor-hosted | Standard LLM rates on input and output | Rough, not calibrated | Low volume, teams already with one provider |
| Your own classifier (e.g. Hugging Face AutoTrain) | Model trained on your labelled data | Either | Training cost, then cheap to run | Real, can be calibrated | One stable, narrow task |
| Safety models (Llama Guard, Granite Guardian) | Classifiers for fixed safety categories | Your own servers | Free licence, you pay for hosting | Per category | Screening prompts and outputs only |
Sources: eesel, DataCamp, Turing Post. Prices as at 24 September 2026.
Each alternative covers part of the ground:
Decision models combine what these offer separately. You write a new question in plain English, with no training data. You get a calibrated confidence score. And you get it in under a second at a tiny fraction of LLM cost. That mix is new, and it is why decision models deserve a place in every fund’s AI design: not as a replacement for LLMs, but as the layer that decides when an LLM, or a person, is needed at all.
Qaike helps hedge funds and private market funds choose, pilot and govern AI tools. Explore the vendors we track in the Qaike AI landscape, or put them side by side with Compare Vendors (requires registration).
This article is based on publicly available information and reflects Qaike’s own analysis and opinions. It is not intended to provide professional advice, endorsement, or a definitive assessment of any vendor or product. While every effort has been made to ensure accuracy, completeness and correctness cannot be guaranteed.
Created by Qaike – Powered by AI.