Ask the smartest AI models in the world about a vulnerability published last month and they do one of two things. They politely decline, or they confidently make something up.
Both score the same.
| Model | Score |
|---|---|
| VCIY sovereign CVE model | 90.86 |
| Claude Opus 5 | 0.0 |
| Gemini | 0.2 |
That is not a typo. We checked, mostly because we assumed it was one.
The test covers vulnerabilities published after every frontier model finished training. Same questions for all three, same rubric, an adversarial judge, and a frozen test instrument hashed so nobody can quietly adjust it afterwards. Including us.
Why zero
A language model knows what was in its training data and nothing after. Asking it about last month’s CVEs is like asking a very well-read historian for this morning’s weather. You will get an answer. It will be beautifully phrased.
The judge only credits factual claims that survive a check against the record. A polite refusal earns nothing. A confident invention earns nothing. Gemini’s 0.2 is roughly what one lucky guess looks like.
None of this is a knock on the frontier models. They are general instruments built to be good at nearly everything, and they are. The vulnerability record is just a uniquely hostile place to be general. It changes every day, and the question a security team actually needs answered on a Tuesday morning is almost always about something that did not exist when the model was trained.
On old news, it is a close race
On older, well-documented vulnerabilities, the picture flips. Our frozen cve-domain-v1 battery has 4,778 items: 3,794 from material our model trained on and 984 it never saw.
| Model | cve-domain-v1 |
|---|---|
| VCIY sovereign CVE model | 88.91 |
| Claude Opus 5 | 88.0 |
| Gemini | 44.7 |
| Untrained base model | 16.9 |
On the historical record, Opus 5 is nearly level with a purpose-built model, which is genuinely impressive and slightly annoying. The gap between 88 and zero is not intelligence. It is the calendar.
Why a later cutoff will not fix it
Every training cutoff is a cliff, and the vulnerability record walks off it daily. Retrain next quarter and you have moved the cliff, not removed it.
What fixes it is retrieval. A model does not need to have memorized last week’s advisory. It needs to pull the right evidence, from the right source, as of the right date, and reason over that. This is the same principle that put our embedding model, Ingot-8B-R3, at the top of MTEB English v2: inference quality is capped by retrieval quality. Hand a capable model the right evidence and it gives a good answer. Hand the best model in the world the wrong evidence and it gives a fluent, confident, wrong one, often with bullet points.
A record that remembers what it used to say
Most vulnerability sources overwrite themselves. When an assessment changes, the old one is gone, and nobody can show you what the record said on the day you made your call.
VCIY keeps two clocks for every fact: when it became true in the world, and when the record learned it. Ask a question as of March 3 and you get what was knowable on March 3. Evidence from after that date cannot appear in the answer. That is enforced by how the record is built, not by asking the model nicely, so the leaked-citation rate is zero.
This matters most at an unglamorous moment: the post-incident review, the audit, the email six months later that begins “Quick question about March.”
It runs where your data lives
The model, the retrieval and the record all run on the customer’s own hardware, with no outbound connection. The most sensitive thing a security team owns is the list of what it runs, and that list never has to leave the building to get an answer.
It is also the only way to get a current answer inside a network that is not allowed to call out. A frontier model behind an air gap is a very expensive time capsule.
What to take from this
If your workflow asks a general model about vulnerabilities, the blind spot is not random. It is exactly the newest material, which is exactly where attackers are working. And the model will not tell you it is blind. It will answer anyway, in complete sentences. That is the dangerous part.
Look for yourself
Every CVE has its own public page at vciy.com. Start with one you know, Log4Shell, CVE-2021-44228, or go back to the beginning with CVE-1999-0025, a buffer overflow old enough to rent a car.
The method, both batteries and every score are published at app.vciy.com/validation.
I wrote the more personal version, including what changed in April 2026 that made all this necessary, on my own blog: NIST Stopped. So I Kept Going.
Jonathan Corners · Founder, Voxell, Inc. vciy.com · voxell.ai