AI startup · Financial due diligence
A due-diligence AI that scored 100% on the Vectara RAG benchmark
Investment teams load a whole data room and ask questions in plain English. Every answer cites the exact pages it came from. I built the platform from scratch as CTO of an AI startup in financial due diligence. It has indexed 10M+ pages and scored 100% on the Vectara RAG benchmark across 3,000 queries.
- pages indexed
- 10M+
- on the Vectara RAG benchmark, 3,000 queries
- 100%
- on SpreadsheetBench, verified by the benchmark team
- 95.5%
- Role
- CTO
- Client
- Name withheld under NDA
The problem
A deal comes with a data room full of PDFs, spreadsheets and call transcripts. Analysts need answers they can defend, so every number has to trace back to the page it came from.
A general chatbot can't promise that. It blends sources, and sometimes it makes things up.
My role
I was the startup's CTO. I built the platform from scratch and wrote most of the ingestion and search code myself. Coding took about 70% of my time. I also ran hiring and owned the infrastructure and the SOC 2 controls.
Constraints
- Every answer needs a citation down to the page.
- Data rooms mix PDFs, charts, spreadsheets and transcripts.
- One firm's documents must never turn up in another firm's answers.
Decisions
Make the page the unit of retrieval
- Context
- The usual approach cuts documents into chunks of text, and a chunk doesn't map cleanly to a page someone can open and check.
- Decision
- Every page is read and indexed on its own, with its file and page number carried all the way through. Gemini Flash-Lite reads every page and flags charts and graphs, and Gemini Flash takes a second pass at only the flagged pages. Each page moves through its own state machine with retries, so one bad page doesn't hold up the rest of the file.
- Consequence
- Search hands the model specific pages with their IDs, never a whole document, so a citation is just metadata that comes along with the page. Each deal's pages live in their own index, so a search only reaches the deals the analyst asked about.
Search three ways, then merge
- Context
- Analysts ask in their own words. The page that answers might use the exact term, like EBITDA, or say the same thing differently.
- Decision
- Each question fans out into several searches written by an LLM. Retrieval runs vector search, BM25 keyword search and HyDE, where the model drafts a likely answer and searches with that, and reciprocal rank fusion merges the results.
- Consequence
- Exact terms and paraphrases both get found, and pages that more than one method finds rank higher.
Check the pages before writing the answer
- Context
- Search always returns something, even when the documents don't hold the answer. A page that only shares keywords with the question is where made-up answers start.
- Decision
- Two checks run before anything is written. First, a fast model reads the retrieved pages, picks the ones that directly answer the question and quotes the exact text that proves each one. Those go to the front of the list. Then the agent reads everything it gathered and drops any page it's sure is off-topic, so it never reaches the answer model or the citations.
- Consequence
- The answer model works from pages both checks have read, and it cites them by ID, so it can't cite a page it never saw.
The result
The platform has indexed 10M+ pages and scored 100% on the Vectara RAG benchmark across 3,000 queries. Its spreadsheet agent scored 95.5% on SpreadsheetBench, verified by the benchmark team. The company also completed a SOC 2 Type II audit, and I owned the controls behind it.
Stack
- TypeScript
- Node.js
- Gemini
- TurboPuffer
- PostgreSQL
- Prisma
- Redis
- Docker