Projects

Things I've built, what I measured, and what didn't work.

DeepTrace

A deep research agent that shows its work.

2026 · open source
  1. Plan
  2. Execute
  3. Report
  4. Verify

I built it from scratch. It plans a research question, searches and reads sources, writes a report where every claim carries a citation, and then a judge model checks each citation against the source it points to. Runs are checkpointed, so a crashed run resumes where it stopped. Follow-up questions reuse the evidence that was already collected.

82 / 82
forced crashes (SIGKILL) recovered, with byte-identical final reports
99.2%
of swapped citations caught by the judge (131 of 132), with 0 false alarms on 129 untouched claims
−19.9%
tokens when each section sees only its routed evidence instead of all of it
What didn't work (yet)

Routing each section to only its own evidence cut tokens, but the share of fully supported claims fell by 5.6 points (95% CI −9.4 to −0.9). Putting the judge in the loop lowered unsupported claims by only 1.1 points, which is within noise (95% CI −5.3 to +4.0). A repair step for partially supported claims is written but not yet evaluated. All numbers come from the eval scripts in the repo, graded by an independent gpt-4o.

PythonasyncioFastAPI + SSELangGraphTavilySQLAlchemy FaissNeo4jReact + TypeScriptDocker

llm-syllogism

Benchmark and code for my AAAI 2026 paper.

2025

Generates 11,000 syllogisms across all 44 classical forms from WordNet hypernym chains, then queries five LLM APIs concurrently with retries, rate-limit backoff and resumable runs.

PythonasyncioWordNetOpenAI-compatible APIs

Literary Writing Agent

Writes new paragraphs in an author's voice.

2026

Takes a story outline, retrieves stylistically close passages from one author's work, and generates paragraphs in that voice. Chunks are cut where the embedding distance between neighboring sentences jumps, which turned 97K lines of prose into 3,000+ retrievable chunks.

LangChainSentence-BERTChromaDBStreamlit