Claims, not checklists
The unit of review is a claim from the paper — "standard errors are clustered by industry" — not a generic linting rule. Every verdict is about whether the code does what the text says.
Evidence you can follow
Verdicts come with pointers like analysis/table3.do#L47 — the exact line that supports or contradicts the claim, so you can check the checker.
Honest about limits
When the package doesn't contain enough to decide, the verdict says uncertain or unverifiable — no false confidence.
benchmarks soonBuilt for real research code
General-purpose LLM tools are tuned on clean software-engineering benchmarks. Dodo Review is built for the code scientists actually write — messier, less tested, and more error-prone. Benchmark numbers will land right here.