The Emerging-Manager Deal Sourcing Playbook: Building Proprietary Deal Flow on Public Data
A 90-day playbook for emerging managers: build a sourcing engine on public engineering data, run detection-triage-outreach in two hours a week, and produce the dated receipts LP diligence actually asks for.
Key Takeaway
Emerging managers cannot out-network platform funds; they out-observe them. This playbook builds the three-part engine, coverage of thesis-sector GitHub orgs, weekly momentum detection, and a scoring triage, on public data in under two hours a week, installs it over 90 days with calibration before outreach, and produces the dated pre-announcement receipts that LP sourcing diligence rewards.
Emerging managers, funds one through three, raise on track record they do not yet have. The pitch to LPs is process: a repeatable sourcing engine, a disciplined filter, and receipts showing the process found real companies early. This guide is the playbook for building that engine on public data instead of headcount.
The emerging-manager sourcing problem#
An emerging manager's constraint is not ideas; it is coverage with zero analysts. You cannot out-network Sequoia, and you cannot out-research a platform fund. What you can do is out-observe: monitor the public exhaust of thousands of startups, engineering activity, hiring, product changes, and arrive at the interesting ones before the network hears about them. The deals-before-databases approach is the core pattern: proprietary observation beats contested referrals.
Building the engine on public data#
The engine has three components. Coverage: a monitored universe, in practice the GitHub organizations of startups in your thesis sectors. Detection: weekly momentum scoring across commits, contributors, and repositories, normalized within sector and stage; the momentum ranking does this across 15 sectors with a free API. Triage: a scoring rubric that converts detections into outreach decisions, templated in the deal flow scoring framework. The whole loop runs in under two hours a week, which is the point: it must survive fundraising season.
What LPs actually ask about sourcing#
In diligence, LPs ask three sourcing questions: how do you see deals others do not, what is your pre-announcement share, and show me the receipts. The third is the differentiator. A dated log of companies identified before their rounds, with the signal snapshot at identification, is an artifact almost no emerging manager has. The scout receipts pattern, verifiable identification before announcement, is directly transplantable to fund diligence.
The 90-day installation plan#
Install in quarters. Days 1-30: pick two sectors, subscribe to the weekly panel, and run detection-only; no outreach, just calibration of what acceleration looks like in your niche. Days 31-60: add triage and a simple pipeline system, begin outreach on the top decile. Days 61-90: publish the cadence internally, log receipts, and cut what you have not used. The failure mode is starting outreach in week one, when every signal looks actionable and the base rates are unknown.
Key takeaways#
Emerging managers win by out-observing, not out-networking. The engine is coverage, detection, triage on public data, installed over 90 days with calibration before outreach. And the LP-diligence artifact that closes is receipts: a dated, verifiable log of companies you identified before their announcements.
Case pattern: the first fundable discovery#
What does success look like in month two. A manager tracking developer-tools sees an org appear in the weekly movers list: commit velocity up 60 percent over its trailing baseline, two new integration repositories, first external contributor. Cross-check adds an engineering job posting posted nine days ago. The manager reaches out referencing the integration work, learns a launch and a round are both in motion, and gets a meeting a month before the round is announced. That single dated receipt, signal snapshot, outreach date, meeting outcome, is worth more in the next LP conversation than a thesis deck, because it demonstrates the engine working end to end.
This is the pattern repeated: detection from public data, cross-check on a second layer, evidence-first outreach. The deals-before-databases guide walks the same loop with more operational detail.
Budget and tooling reality#
The engine runs on free tiers: the public signals API and MCP server for coverage, job boards and changelogs for cross-checks, a spreadsheet for the pipeline until volume justifies a CRM. The paid stack, contact enrichment, panel history, becomes worth it only when outbound volume makes manual lookups the bottleneck; the free VC tools answer and the deal flow tools comparison price those upgrade points honestly. Until then, the constraint is never data cost; it is cadence.
Failure modes for new managers#
Three failures recur. Starting outreach before calibration: week-one signals all look actionable, and burned outreach cannot be un-burned. Monitoring too wide: fifteen sectors with no thesis is a dashboard, not a strategy; two sectors read deeply beats fifteen scanned. And stopping the log: the receipts file is the compounding asset, and it only compounds if every outreach, hit or miss, gets an entry. The managers who keep the log through the boring weeks are the ones holding LP-grade evidence at month eighteen.