A research assistant that refuses to answer
- Client
- Fathom Capital
- Sector
- Financial services
- Year
- 2024
- Services
- AI Systems
94%
answers with correct citations
0.4%
unsupported claim rate
6.5hr
analyst time saved per week
Challenge
Analysts wanted natural-language search over fifteen years of internal notes, filings and call transcripts. In a regulated setting a confident wrong answer is worse than no answer, and the first prototype the team had built was fluent, plausible and wrong often enough to be unusable.
Approach
In order, including the steps that changed the plan.
- 01
Score refusals as successes
The eval set includes questions the corpus genuinely cannot answer. A confident answer to those scores zero. Optimising against that changed the system design more than any prompt did.
- 02
Cite or do not claim
Every sentence in a response is checked against its retrieved sources. Unsupported sentences are removed before the response is shown, and if too little survives, the system says so.
- 03
Make provenance one click away
Citations link to the exact passage and document version. Analysts verify in seconds, which is what turned occasional users into daily ones.
Outcome
Unsupported claims sit at 0.4% against a 5% acceptance threshold. Adoption climbed steadily through the year, which the team attributes to the system being reliably silent when it should be.
Built with
- Claude
- pgvector
- Braintrust
- Next.js
- Postgres
Tell us what you are building.
A system you want built, a model that has to survive real traffic, or a process that should have been automated a year ago. The first conversation costs nothing, and occasionally ends with us telling you not to build it.
Reply within one working day
A person, not an autoresponder.
A 30-minute call, no deck
We ask about constraints, not budget.
A written view within a week
Including the case for not proceeding.