DeepTrace
A deep research agent that shows its work.
- Plan
- Execute
- Report
- Verify
I built it from scratch. It plans a research question, searches and reads sources, writes a report where every claim carries a citation, and then a judge model checks each citation against the source it points to. Runs are checkpointed, so a crashed run resumes where it stopped. Follow-up questions reuse the evidence that was already collected.
What didn't work (yet)
Routing each section to only its own evidence cut tokens, but the share of fully supported claims fell by 5.6 points (95% CI −9.4 to −0.9). Putting the judge in the loop lowered unsupported claims by only 1.1 points, which is within noise (95% CI −5.3 to +4.0). A repair step for partially supported claims is written but not yet evaluated. All numbers come from the eval scripts in the repo, graded by an independent gpt-4o.