Verify
Check an AI agent's research, claim by claim
Paste a research memo an AI wrote. Litmuz breaks it into individual claims, checks each citation against the primary literature, and flags anything fabricated, unsupported, or unsafe. It triages and flags; it never certifies on its own.
- 1 Every claim is checked against PubMed and PMC.
- 2 A red flag means a fabricated or contradicted citation; yellow routes to human review.
- 3 Safety-critical claims can never auto-pass.
Why Litmuz
AI writes research faster than anyone can check it
AI research agents are fluent, fast, and confidently wrong in ways that are expensive in the life sciences. They cite papers that do not exist, cite real papers that do not support the claim, and state a dose or a target with the same certainty as a fact. Reading the prose, you cannot tell a grounded claim from a fabricated one. Litmuz makes that difference visible, one claim at a time.
Litmuz turns each of these into a verdict you can act on.
How it works
Paste a memo from a research agent. With the literature criterion on, Litmuz runs the memo through four stages and returns a verdict for every claim, with the evidence attached. Two of the stages are rule-based and run with no model in the loop. The stage that decides whether evidence supports a claim uses Claude, and it has to quote the exact sentence it relied on. Genomic evidence takes a different path: no citation registries, no retrieval, no model, and the same safety gate and traffic lights at the end.
- Agent memoinput
- Decompose into claimsClaude
- Check citationsrule-based
- Retrieve + judge evidenceClaude
- Safety gate + severityrule-based
- Per-claim verdict
Rule-based, no model Claude reads the evidence
- 1
Decompose the memo into atomic claims
Claude splits the memo into single checkable propositions and attaches each citation to the claim it is meant to support. Nothing is verified at the paragraph level, because a paragraph can be half right.
- 2
Resolve every citation deterministically
Each identifier is resolved by rule against PubMed, PMC, and Crossref, with no language model involved. A fabricated identifier, a retracted source, or metadata that does not match the claim is caught here, and cannot be argued away downstream.
- 3
Retrieve the evidence and judge whether it supports the claim
Litmuz pulls the passages from the resolved source, and Claude assesses whether that evidence actually entails the claim. A supported verdict must quote the exact sentence it relied on; when the sentence is not there, the claim does not pass.
- 4
Apply the safety gate and return a per-claim verdict
A separate deterministic lexical check re-derives whether a claim is safety-critical (a dose, a dosing regimen, a clinical indication, a molecular target), and those claims are always routed to a human, however confident the model is. Every claim ends on green, yellow, or red, with a diagnostic code and a link to the citation status and evidence behind it.
What the verdicts mean
Every claim gets one of three traffic lights, plus a diagnostic code (D1 to D5) and a link to its citation status and evidence. Green is not the only good outcome. Yellow and red are first-class results: they tell you where the memo is not standing on evidence.
Grounded (green)
The evidence supports the claim and the citation resolves cleanly against PubMed, PMC or Crossref, or (in genomic mode) the region matches the reference. In literature mode Claude has to quote the exact supporting sentence, so you can check the call yourself.
Needs review (yellow)
The claim is unverifiable, under-supported, or safety-critical, so it goes to a human in the review queue. A safety-critical claim (a dose, a dosing regimen, a clinical indication, a molecular target) always lands here, even when the evidence looks supportive and the model is confident. A retracted source or an expression of concern can never come back green.
Flagged (red)
The evidence contradicts the claim, or the citation is fabricated. Citation resolution is deterministic, so a made-up identifier is caught by rule, not by judgment.
Two ways to verify
Pick the criterion that fits the memo you are checking. Each run is scored against one of them. Every verdict still carries its diagnostic code and links back to the evidence that produced it.
Literature
Citations are resolved deterministically against PubMed, PMC and Crossref, with no model in the loop. A fabricated identifier, a retracted source, or metadata that does not match the claim is caught by rule, not by judgment. Only then are evidence passages retrieved for Claude to judge entailment, and it must quote the exact supporting sentence.
Genomic evidence
Genomic claims are checked against Gladstone Institutes reference data: the Pollard lab's Human Accelerated Regions and Zoonomia mammalian conservation. The check is deterministic, with no LLM involved at any step, so a real HAR asserted truly comes back green and a real HAR denied comes back red. The reference is a curated subset of well-characterized HARs rather than the full catalogue, so a region it does not cover returns an honest needs-review and is routed to a person, never a silent pass.
What it checks against
The sources are public. The prompts, thresholds and rubric that combine them are not.
Literature
Citations resolved deterministically, then evidence retrieved
- PubMedNCBI E-utilities (esummary, efetch) for identifiers, metadata and abstracts
- PMCPubMed Central open-access full text (BioC) when a paper is available
- CrossrefDOI resolution and publication metadata
- Retraction statusRetracted papers and expressions of concern flagged from the source record
Genomic evidence
Gladstone Institutes reference data, checked deterministically
- Human Accelerated RegionsCurated, well-characterized HARs from the Pollard lab
- ZoonomiaMammalian conservation across 240 placental mammals
Language model
Used only where reading is required, never to decide a pass
- Anthropic ClaudeDecomposing the memo into claims and judging whether retrieved evidence entails a claim
- Every gate is rule-basedCitation resolution, retraction, the safety gate and the traffic-light mapping run with no model
Who it is for
Litmuz is built for the people who have to stand behind what an AI agent wrote, and for the agents that would rather not hand over weak work in the first place.
The computational biologist reviewing an agent's output
An agent returns a fluent memo where every sentence sounds equally confident and every citation looks equally real. Litmuz splits it into atomic claims and marks each one green, yellow or red, so you can start with the claims that are contradicted or unsupported instead of reading every sentence at the same level of suspicion. Green claims still carry the quoted sentence they rest on, so you can check any call yourself.
Teams running AI agents in drug-discovery and genomics workflows
When a target or a dosing claim gets questioned three months later, "the model said so" is not an answer. Litmuz stores every verdict with the evidence behind it, and every human decision is written to an append-only audit trail with its rationale, reviewer and timestamp. Safety-critical claims never auto-pass into that record without a person seeing them.
The AI agent itself, over MCP
Any Claude-powered agent can call Litmuz as a tool and get its own claims back with per-claim verdicts and diagnostic codes. It can then drop the fabricated citation or soften the overreaching sentence before a human ever reads the draft.
Frequently asked questions
How the verification works, what the verdicts mean, and where the limits are.
What is Litmuz?
Litmuz is a claim-level verification layer for life-sciences research agents. You paste a memo written by an AI research agent, and Litmuz decomposes it into atomic claims and returns a traffic-light verdict for each claim, with the citation status and the supporting evidence attached. It is open source and runs at litmuz.co. It triages and flags; it never certifies on its own.
How does Litmuz verify a claim in an AI-generated research memo?
It splits the memo into atomic claims and checks each one separately. Citations are resolved deterministically against PubMed, PMC and Crossref, with no model involved. Evidence passages are then retrieved and Claude judges whether the passage entails the claim, and it has to quote the exact supporting sentence. Every claim comes back with a traffic light, a diagnostic code (D1 to D5), and links to the citation and evidence behind it. The full report exports as CSV.
What do the green, yellow and red verdicts mean?
Green means grounded: the evidence supports the claim and the citation resolves cleanly. Yellow means needs review: the claim is unverifiable, under-supported, or safety-critical, so it is routed to a human, and a retracted source or an expression of concern can never come back green. Red means flagged: the evidence contradicts the claim, or the citation is fabricated. Yellow and red are first-class outcomes, not failures. The core principle is honest negatives, so an unsupported claim is never dressed up as a pass.
Is the Litmuz verdict just an LLM guessing?
Not for the parts that can be settled by rule. Citation resolution and genomic checking are deterministic and rule-based, with no model in the loop, so the same input gives the same result every time. Claude is used to split the memo into claims, to categorize them, and to judge whether a retrieved passage entails the claim. Only the judge produces the verdict, and it must quote the exact sentence it relied on so you can check its work. Rules also cap the model, so a fabricated citation is red no matter how confident the judge was, and a retracted source can never be green.
How does Litmuz catch fake or hallucinated citations?
Every identifier in the memo (PMID, DOI, PMCID) is resolved against PubMed, PMC and Crossref by rule. A well-formed identifier that is authoritatively absent is marked fabricated, which forces the claim to red. If the identifier resolves but the title or authors do not match what the memo attributed to it, that is a metadata mismatch, and it is recorded on the citation status and downgrades the claim, so it can never come back as a clean D1 pass. Retracted sources and expressions of concern are detected separately and can never be green.
What happens to safety-critical claims like a drug dose?
They can never auto-pass. A dose, a dosing regimen, a clinical indication, or a molecular target is always routed to a human, no matter how confident the model is. Safety-criticality is re-derived by an independent deterministic lexical check, not just taken from the model's own label, so a claim the model mislabeled still gets caught. The best outcome a safety-critical claim can show on its own is yellow.
What is genomic verification mode in Litmuz?
It is the second verification criterion you can pick, instead of literature. Genomic claims are checked deterministically, with no LLM at all, against Gladstone Institutes reference data: the Pollard lab's Human Accelerated Regions and Zoonomia mammalian conservation. A real HAR asserted truly comes back green, a real HAR denied comes back red, and a region that is absent from the curated reference is an honest yellow rather than a silent pass. The reference is a curated subset of well-characterized HARs, not the full catalogue of roughly 3,100.
What is a Human Accelerated Region (HAR)?
A HAR is a stretch of genome that stayed nearly unchanged across mammals for millions of years and then changed rapidly in the human lineage. Several sit near genes involved in brain development, which is why they get cited often in evolutionary and neurodevelopmental claims. Litmuz uses a curated set of four well-characterized HARs (HAR1, HACNS1, HARE5 and 2xHAR.170) as deterministic ground truth for genomic claims. If a claim names a region outside that curated set, Litmuz says it cannot confirm rather than guessing.
Can a human override a Litmuz verdict, and is the override recorded?
Yes to both. Any claim routed for review, yellow or red, lands in a review queue where a reviewer can confirm the machine verdict or override it, and an override requires a written rationale. The decision is stored with who made it and when, in an append-only audit trail, and the claim's final traffic light is re-derived from the human decision. Nothing is silently overwritten, so the original machine verdict stays visible next to the human one.
Can my AI agent call Litmuz directly?
Yes. Litmuz ships an MCP server, so any Claude-powered agent can call verification as a tool over the Model Context Protocol. An agent can submit a memo, poll the job, and pull back per-claim provenance. MCP calls run the same pipeline as the web app, with the same deterministic citation checks and the same safety gate. Sources an agent supplies are additive evidence, and they cannot switch off the deterministic citation check or the safety gate.
How much does Litmuz cost?
The free tier covers 2 verifications per week. Pro covers 100 per week. Quotas run on the UTC calendar week and reset Monday. The code is open source if you would rather run it yourself.
Does Litmuz replace peer review?
No. Litmuz triages and flags; it never certifies on its own. It catches fabricated citations, contradicted claims and under-supported claims fast, and routes anything it cannot settle to a person. It is not a clinical, diagnostic, or regulatory tool, and it is not a substitute for peer review or expert judgment.
Building your own verification or evals?
Litmuz is one instance of a general problem: keeping AI honest in a domain where being confidently wrong is expensive. If your team is designing claim-level checking, honest-negative rubrics, evaluation harnesses, or domain-grounded verification, I am happy to help.
Get in touchCheck a memo before you trust it
Paste an agent memo above, or read exactly how a verdict is reached.