Jev vs LLMs, closed and open
A government RFP hides a few hundred must-do rules in a hundred pages of legal text. Miss one and your bid can be tossed. I marked every rule in three real RFPs by hand, then asked five AI models to find them. The winner can't write a sentence. It reads and picks, for two cents a document.
| Model | Kind | F1 | Per RFP |
|---|---|---|---|
| Jev 1.13 TypeSafe AI | Decision model | 0.64 | $0.02 |
| Kimi K2.6, fine-tuned | Open | 0.57 | $3.30 |
| Qwen3.5-4B, fine-tuned | Open | 0.55 | $0.22 |
| Inkling, fine-tuned | Open | 0.53 | ~$1.20 |
| Claude Opus 4.8, few-shot | Closed | 0.52 | $4.70 |
Higher is better. 1.0 would mean every rule found and nothing wrongly flagged.
What is Jev
Most AI models you've heard of are writers. Ask them something and they compose an answer, word by word, like a consultant billing by the hour. Jev is not a writer. It's a sorter.
Picture the mailroom. A letter comes in and the clerk doesn't write back. They glance at it and drop it in a slot: bills, junk, legal, the boss. Fast, cheap, and nobody gets a five-paragraph memo about the envelope.
That's Jev. I hand it one sentence from an RFP and a few fixed slots: is this a rule, what kind, can it get the bid thrown out. It picks a slot and says how sure it is, in about a fifth of a second. It charges for reading, not talking, because it doesn't talk.
The catch: the clerk can't tell you why. And when Jev says it's 70% sure, it's right about one time in five. Great at sorting, not so great at judging itself. Like most of us.
Or open a past run
The three scored RFPs, plus the enrollment-broker RFP from my July run. That one has no answer key, so it gets a bid read but no score.
Jev is reading
What was the job
The job
When a state agency buys something, it publishes an RFP: a long document listing everything a vendor must do, have, or send. Firms pay people to comb through these by hand. I did it myself for three Texas RFPs and marked all 438 rules. That list is the answer key.
The score
Each model read the same three documents and listed the rules it found. The score goes up for every real rule it catches and down for every sentence it flags that isn't one. Nobody came close to 1.0. That is the honest state of this job today.
What happened
Jev caught 354 of the 438 rules, far more than anyone else, and won on all three RFPs. It also over-flags: about half of what it picks isn't a rule. The LLMs were pickier but missed a lot more. For a bid team, a missed rule costs more than an extra line to cross off.
The catch
I built the answer key one sentence at a time, and Jev picks one sentence at a time, so the test suits it. And don't trust its confidence in the middle: when it said 50 to 80% sure, it was right 18% of the time.
Model by model
- Jev 1.13 0.64 · $0.02 a document
- Doesn't write at all. It reads each sentence and answers a fixed question: is this a rule, and what kind? Found 81% of the rules. The cheapest by a mile.
- Kimi K2.6, fine-tuned 0.57 · $3.30
- A big open model I trained on 11 practice RFPs. The best of the LLMs. Found about half the rules, and most of what it flagged was real.
- Qwen3.5-4B, fine-tuned 0.55 · $0.22
- The small one. After the same training it nearly matched Kimi at a fifteenth of the price. It also picked up a habit of repeating one line hundreds of times, which I had to catch and rerun.
- Inkling, fine-tuned 0.53 · ~$1.20
- Trained the same way, reading each RFP in pieces. Uneven: on the biggest RFP it found 73% of the rules, the best any LLM did anywhere, then slipped on the other two.
- Claude Opus 4.8, few-shot 0.52 · $4.70
- No training, just a few worked examples in the prompt. The most careful: on one RFP, 91% of what it flagged was right. But it found only 4 in 10 rules, and one missed rule can sink a bid. The most expensive.
- The open models, untrained 0.00 to 0.48
- Before training they mostly talked to themselves. They wrote pages of notes about each rule and ran out of room before giving an answer.
The rules
Jev answers. Code decides.
A missing qualification is a no
Years in business, past contracts, a certification you don't hold. One gap at 70% confidence ends it.
Paperwork is a checklist
You can order a certificate by Friday. You can't order five years of history.
Coverage decides the rest
Half the required work covered is a bid. Less is a closer look.
Where Jev breaks
Its confidence. Answers it gave 50 to 80% were right 18% of the time.
Reads 20 times more. Costs 2 cents.
| Model | Tokens read | Tokens written | Calls |
|---|---|---|---|
| Jev 1.13 | 1.35M | 0 | 1,311 |
| Kimi K2.6, fine-tuned | ~66K + prompt | ~36K | 9 |
| Claude Opus 4.8, few-shot | ~66K + prompt | ~21K + thinking | 3 |
Reading runs in parallel. Writing goes one token at a time, and that is where an LLM's bill goes.
The answer can't break
A pick from a fixed list. No JSON to parse, no loop repeating one row hundreds of times.
The bill is known up front
It follows document length, not how much a model decides to say.
Failures stay small
A bad call loses one sentence, not a chunk of the RFP.
What you give up
1,311 calls to manage, options you must write in advance, and no reasons.
Same pattern, other jobs
Security questionnaires, support ticket routing, invoice coding, content moderation, the routing step inside an agent.