Most tools that answer a question this way are asking a model what it thinks and dressing the reply in citations. This one gathers evidence, scores it by a fixed rule, and shows you the arithmetic. The difference is not a matter of degree, and it is worth being able to check.
What the machines do, and what they do not
Epistry uses language models, and it would be strange to pretend otherwise. They do four jobs: work out what testable claim is inside your question, search for and summarise sources, code each piece of evidence (what kind of study is this, does it support or oppose the claim, how big was the sample), and argue the case from several sides so you can read the disagreement.
What they never do is decide. The verdict is computed from the coded evidence by a fixed formula with no model in it. Ask the same question twice and the arithmetic returns the same answer; two people asking on opposite sides of an argument get the same number. That is the whole design, and everything below is a consequence of it.
The debate you can read on a report — the advocate, the skeptic, the methodologist, the arbiter who sums them up — is genuinely useful and completely powerless. None of it moves the score. It is there because a verdict you cannot argue with is not worth much, not because it votes.
The short version
Models read. Arithmetic decides. The two are kept apart on purpose, and a test in the build fails if they ever touch.
The sources are real, and checked
Evidence is retrieved, not remembered. Depending on the claim, that means academic databases with real citation graphs and real retraction flags, news archives, or a neural search over the open web — and for scientific questions, the reference lists of the papers that came back, so the foundational work a topical search would miss is pulled in too.
Anything the model reports that was not in the real results is dropped before scoring. If retrieval fails, the report says the evidence base was incomplete rather than quietly scoring a thin pool — and that run does not cost you a token.
Opposition is searched for deliberately, not just noticed if it turns up. A separate pass goes looking for the case against, and what it finds is scored alongside everything else. When that search comes back empty, the report distinguishes “we looked and found no credible opposition” from “we did not look”. Those are very different statements and most tools conflate them.
Why it matters
A language model asked for citations will happily invent them. Every reference here came back from a search that actually ran.
Evidence is weighed, not counted
Counting sources rewards whoever published most. So each piece of evidence starts from what kind of evidence it is — a systematic review of many trials outranks a single trial, which outranks an expert’s opinion, which outranks an anecdote — and is then adjusted for the things that make evidence better or worse.
It is discounted for failing to replicate, for age where age matters, for small samples, for retraction, and for reasoning that commits a recognised fallacy. Scientific evidence picks up the GRADE downgrades on top: indirectness (a mouse study is not a human outcome), imprecision (an effect can be statistically real and practically trivial), and risk of bias in how the study was run.
Then the count itself is corrected. Ten outlets running the same wire story are one story, and a citation graph is used to work out which sources are genuinely independent and which are echoes of each other. A claim supported by five independent research groups is on firmer ground than one supported by fifty articles that all trace back to a single press release, and the score reflects that.
The lineage
The downgrade structure follows GRADE, the framework clinical guideline panels use to rate certainty in evidence.
Two axes, because "how sure" and "which way" are different questions
The weighed evidence produces two independent numbers. The balance between supporting and opposing weight sets the direction. The number of independent sources behind it sets the confidence.
Keeping them apart is what lets the system say something most tools cannot: the evidence points this way, but there is not enough of it to rely on yet. A lopsided result resting on two related papers is reported as leaning, not settled, no matter how lopsided it is.
It also means “not enough evidence” is a real answer rather than a failure. It is never rendered as though the claim were false — an absence of evidence and a refutation are different findings, and conflating them is how confident nonsense gets manufactured in both directions.
What you see
Direction is the hue on a report; confidence is the width of the fan. A colour never has to carry both.
It commits to what would change its mind — first
Before any evidence is gathered, the system writes down what would have to be true for the claim to be wrong. Only then does it search. At the end, each condition is checked against what was actually found, and the result is shown to you — including the uncomfortable case where a condition was met.
The ordering is the entire point. Falsification conditions written after the evidence is in are a formality; written before, they are a commitment. It is the same discipline a pre-registered study uses, for the same reason.
Order matters
The conditions are written before retrieval, so they cannot be shaped to fit the answer that turned up.
What we are not claiming
The verdict is only as good as what was retrievable. On a story breaking today, the evidence base is thin and the report will say so rather than manufacture confidence from it.
Coding evidence is a judgement, and a model makes it. The structure constrains that judgement hard — a source’s claimed strength is capped by what the source actually is, so a blog post cannot enter the arithmetic as a clinical trial — but it does not eliminate it.
And the thresholds themselves are set from the evidence-appraisal literature rather than fitted to a test set. Tuning them against the small set of claims we can check would make the scores look better on those claims and mean nothing anywhere else. They stay where the published frameworks put them until there is a corpus large enough to say otherwise honestly.
Every one of those limits is on the report itself, next to the verdict, rather than in the small print here.
Honest limits
Calibration of the numeric thresholds is deliberately paused rather than tuned against a small test set.
Try it on something you disagree with
A method is easiest to judge on a claim you already have a view about. Bring one, and read what comes back — including the part where it tells you how it could be wrong.