Product Requirements Document

One Agentic — VC Deep Research Agent

Automated company research pipeline: from pitch deck upload to structured due diligence report in under 20 minutes, with parallel AI agents, conflict detection, and thesis fit scoring.

Status
Draft
Date
13 May 2026
Owner
Product · Engineering
Pipeline stages
02 → 03 → 04

Problem & Goal

North star: A VC partner reads the hybrid report before any meeting and has already seen everything findable before the founder walks in — in under 20 minutes, from any input.

VC analysts spend 2–10 hours on a first-pass diligence for a single company. Most of that time is gathering publicly available information fragmented across dozens of sources — LinkedIn, corporate registries, news archives, SEC filings, GitHub. The result: inconsistent depth across deals, deals skipped due to backlog, and conflicts between founder claims and public records that only surface late — if at all.

What we're solving

  • Manual research takes 2–10 hours per company
  • Inconsistent depth — depends on who ran it and when
  • Claim validation happens in the meeting, not before
  • Deals skipped due to analyst backlog, not lack of merit

Who uses this

  • VC Analyst — runs research, reviews findings
  • VC Partner / GP — reads scorecard and conflict panel
  • Accelerator Manager — same workflow, higher volume

User Stories

1
As a VC analyst, I upload a pitch deck PDF and receive a structured report within 20 minutes so I can prepare for a screening call without manual research.
2
As a VC analyst, I enter a company name or URL and get a report even when I don't have a deck, so I can quickly research inbound cold emails.
3
As a VC partner, I open the scorecard and immediately see thesis fit score, conflict count, and overall confidence — so I can decide in 30 seconds whether to read further.
4
As a VC analyst, I see exactly which deck claims could not be validated and which are directly contradicted, so I know what to probe in the call.
5
As a VC analyst, I see the research run transparently — which agents completed, which found no data, which sources were unavailable — so I can calibrate my trust in the output.
6
As a VC analyst, I am never charged credits for a run that failed before producing a report.

F1 — Input Ingestion

Required

Four input types are supported. All produce the same company profile and claimed metrics before any research begins.

Input typeWhat happensClaim validation?
Pitch deck PDFDocument uploaded; AI extracts company profile and all claimed metrics with slide referencesYes — slide number stored per claim
PPT / generic PDFSame process, lower confidence; falls back to name-only if document is unreadableYes — lower reliability
Company nameCompany identified and resolved from public registries — no document requiredNo — claim validation skipped
Company URLAI reads the company's public web presence and extracts profile informationPartial — from marketing copy

Acceptance criteria

  • All four input types produce a company profile before any research begins
  • Name collisions prompt user to confirm which entity before research starts
  • Unreadable files fall back gracefully — user is notified, no silent failure
  • Low-confidence extraction (fewer than 3 key fields) shows "Limited source material" banner on the report

F2 — Parallel Research Pipeline

Required

Eight specialist research dimensions run simultaneously. Five are required for every run; three are best-effort and may return no data for early-stage companies where public information is limited.

Required Company Overview
Corporate registration, incorporation history, public filings, web presence
Required Founders & Team
Founder backgrounds, career history, developer activity, sanctions screening
Required Web & Digital Presence
Website traffic trends, domain age, web archive history
Required News & Media
Press coverage, news sentiment, article volume over time
Required Claim Validation
Runs after all other dimensions complete. Produces a verdict for every claim in scope across all three claim sources.
Optional Market Analysis
Market size signals, sector trends, competitive landscape — best-effort for early-stage
Optional Legal & Regulatory
Patent filings, public litigation signals, regulatory and sanctions flags
Optional Financial Signals
Public financial filings, funding announcements — most pre-seed companies return no data
Claim sourcing — three tiers, always active regardless of input type:
  • Extracted claims — pulled from pitch deck, company website, or marketing copy (e.g. "$10M ARR", "deep tech", "fastest growing in MENA")
  • Thesis-driven claims — derived from the analyst's saved thesis (e.g. "must have traction", "must operate in fintech"). Always present regardless of input type; defaults to a generic early-stage rubric until a thesis is configured.
  • Analyst-supplied claims — when the system finds fewer than 3 extractable claims, the analyst is prompted to enter claims manually before the run begins.
For name or URL inputs where no pitch deck is available, claims are sourced from the company's public web presence and the analyst's thesis. Claim Validation is never skipped.

Acceptance criteria

  • All required dimensions run simultaneously; total research time under 20 minutes
  • Optional dimensions with no data return cleanly — no error shown to user, absence acknowledged
  • Claim Validation runs only after all other dimensions have completed
  • Claim Validation produces a verdict for every claim in scope — verdict types: confirmed, contradicted, unvalidatable, or awaiting founder response
  • If fewer than 3 claims are found before run start, the analyst is prompted to supply claims manually; the run does not proceed until at least one claim exists
  • Each dimension result includes a status, confidence score, prose summary, findings, and sources consulted

F3 — Conflict Detection

Required

The system cross-references all research findings and surfaces three types of conflicts. This is the core differentiating capability — not data aggregation, but reasoning about inconsistencies across independent sources.

TypeDefinitionExample
Implausible Absence A claimed metric has zero corroborating data points across all sources Deck claims "$10M ARR" — no revenue signal found in any source
Cross-Source Conflict Two agents report incompatible facts about the same entity Founders agent finds 2 co-founders; Company agent finds 3 named in filings
Manufactured Signal A strongly positive signal appears in only one source with no corroboration One article claims "fastest growing SaaS in MENA" — no other mention found

Acceptance criteria

  • Each conflict includes: type, severity (high / medium / low), plain-language description, source A vs source B, and which dimensions surfaced it
  • Conflicts are sorted by severity in the report (high first)
  • Conflict detection runs even if optional dimensions did not complete
  • Zero conflicts → Conflict Panel is hidden entirely, not shown as empty

F4 — Hybrid Report Output

Required

The report is a web page with four zones. It streams progressively — the scorecard is visible within ~30 seconds while agents are still running.

1
Scorecard — streams first, above fold
Visible within 30s
  • Thesis Fit Score — 0–100, scored against the fund's saved thesis
  • Overall Confidence — 0–1, averaged across completed dimensions
  • Conflict count — with severity breakdown (e.g. "1 high · 2 medium")
  • Research metadata — total time, source count, agents completed
2
Conflict Panel — only shown when conflicts > 0
  • One card per conflict, sorted high → medium → low severity
  • Each card: type badge · plain-language description · source A vs source B · agents involved
3
8 Expandable Sections — one per research dimension
  • Collapsed by default; header shows confidence badge + conflict count + chart indicator
  • Expanded: prose summary · data points table · sources list
  • Dimensions with no data show "Insufficient public data at this stage" — never empty tables or broken UI
4
Sources Footer
  • Deduplicated list of all sources consulted across all agents
  • Unavailable sources explicitly noted: "[Source] — unavailable during this run"
Live progress: The report streams progressively — each research dimension populates its section as it completes. The analyst sees real progress rather than a spinner, and can read completed sections before the full run finishes.

Acceptance criteria

  • Scorecard is visible within 60 seconds of run start, before all dimensions complete
  • Dimensions with no data never show empty tables or missing elements
  • All sources cited in prose summaries and in the sources footer
  • Report saved and accessible via shareable URL — survives browser refresh

F5 — Error Handling & Credit Policy

Required
Credit rule: Credits are deducted only on successful report generation. Failed runs consume zero credits. Success is strictly defined: Company Overview, Founders & Team, and Claim Validation must all complete. Claim Validation is considered complete when it has produced a process-based verdict (confirmed / contradicted / unvalidatable / awaiting founder) for every claim in scope — regardless of verdict type or outcome.
FailureSystem behaviourUser sees
Required agent timeout (>3 min) Run continues; that dimension marked low-confidence "Research timed out — partial data only" banner on that section
Optional agent — no data Run continues; dimension excluded from conflict analysis "Insufficient public data at this stage" — no confidence score shown
Data source temporarily unavailable Automatically retried; if still unavailable, marked as skipped Sources footer: "[Source] — unavailable during this run"
Run fails to start (input error or system fault) Run aborted; no credits consumed; user can retry "Research could not start — [reason]. No credits consumed. Retry available."
All required agents fail Run aborts after reducer detects no required agent completed Same as above — zero credits consumed

Confidence score rules

LevelWhen shown
HighTwo or more independent sources agree · No contradictions · All key fields populated
MediumSingle source · Or two sources with minor discrepancy · Some fields missing
LowConflicting data · Dimension timed out · Low extraction confidence
Not shownOptional dimension returned no data — absence acknowledged, score omitted

Success Metrics

<20 min
P95 research time (5 required agents)
<60 s
Scorecard visible from run start
<3%
Failed runs (no report produced)
>97%
Required agent completion rate
>80%
Conflict precision — validated by analyst in pilot
0%
Credits charged on failed run

Constraints & Non-Goals

LinkedIn: No public API exists — founder research relies on publicly indexed content only, not structured profile data. Founder history depth is limited compared to a manual LinkedIn search.

Known constraints

  • Web traffic data requires a commercial web analytics subscription. Until active, the Digital Presence dimension falls back to free-tier signals (domain history, web archive snapshots).
  • Certain government-held ownership records are not publicly accessible and are excluded from all research.
  • Financial Signals dimension is expected to return no data for most pre-seed companies — this is expected behaviour, not a gap.
  • Thesis fit score defaults to a generic early-stage VC rubric until Stage 01 (Thesis Capture) is complete.

Non-goals (this PRD)

  • Stage 01 Thesis Capture — separate system
  • Stage 05–06 Human Review and Follow-up Research
  • Auth and multi-tenancy — platform layer
  • Pricing / credit consumption UI — billing system
  • Frontend application shell (nav, login, deal list)

Open Questions

4 unresolved
Q1
Web analytics subscription at launch. Full web traffic data (monthly visitors, traffic trends, traffic sources) requires a commercial subscription — likely SimilarWeb or SEMrush. Should this be included in the launch cost, or should the Digital Presence dimension launch with free-tier signals only (domain age, web archive snapshots) and the subscription added post-launch?
Q2
Thesis fit scoring without a saved thesis. Until a fund configures their thesis via Stage 01 (Thesis Capture — see dedicated PRD), the score falls back to a default rubric built from pre-defined primitives (stage, sector, geography, traction signals, team signals, business model). The open question is what that default primitive set should be, and whether the analyst must confirm it before a run begins or the system silently applies it. See Thesis Capture PRD for full thesis definition requirements.
Q3
Report diff across multiple runs on the same company. When an analyst runs research on the same company more than once (e.g. after a follow-up meeting or a funding announcement), should the UI surface a diff view highlighting what changed between runs — new conflicts, resolved claims, updated signals? If so, is this a display layer feature or does it require the pipeline to produce structured deltas?
Q5
Founder email mechanic. A planned extension where the system emails the founder with specific questions when a claim cannot be validated from public sources, then holds the claim in "awaiting founder" state until a response is received and validated. Not yet decided whether this belongs in this PRD or a dedicated follow-up PRD.