Free VC Data Sources: The Complete Guide for Early-Stage Investors
A complete map of free startup data for investors: the coverage/verification/monitoring/tooling layers, what free data honestly cannot do, the two-hour weekly workflow built on it, and the trigger points for paying.
Key Takeaway
Most data an early-stage investor needs is free. This guide layers the free stack, GitHub and public signals APIs for coverage, Crunchbase free tier and Form D filings for verification, momentum rankings and job boards for monitoring, and defines the honest limits (open-source over-representation, thin history, round-data lag) plus the trigger points at which paying for enrichment beats free.
The expensive platforms are good at what they do, but a majority of the data an early-stage investor needs is free. This guide maps the free tier of the startup-data world: what each source actually covers, what it cannot do, and how to combine them into a working stack that costs nothing and beats a naive expensive stack at sourcing timing.
The free stack, by layer#
Coverage layer, who exists: GitHub itself, the org pages of every startup that builds in public, plus the public signals API at /api/signals.json which aggregates hundreds of startup orgs across 15 sectors weekly. Verification layer, who raised: Crunchbase's free tier, announced rounds on TechCrunch and DealRoom, and SEC Form D filings for US deals. Monitoring layer, what changed: GitHub watch lists, the weekly momentum ranking, and job boards for hiring signals. Tooling layer: the best startup signal tools list separates the free-core tools from paid ones honestly.
What free data cannot do#
Honest limits. Coverage: free sources over-represent open-source and developer-facing companies and under-represent stealth and enterprise-internal builds. History: free tiers rarely give panel history, so slope calculations need your own logging or a public archive. Freshness on rounds: funding databases lag announcements by days to weeks. None of these limits touch the core job, finding companies before the round, where free engineering data actually leads the paid databases.
Combining them into a workflow#
The free stack as a workflow: Monday, pull the weekly movers from the momentum panel, note sector and stage. Wednesday, cross-check the top decile against job boards and product surfaces. Friday, outreach to the two or three companies that cleared triage, logging the receipt for each. Two hours total. The weekly sourcing workflow template formalizes it, and the free VC tools for emerging managers answer covers the tooling layer.
When to graduate to paid#
Pay when a specific job breaks, not for coverage FOMO. The trigger points: investor-contact enrichment when outbound volume makes manual lookup uneconomic, panel history when slope analysis becomes central, and relationship intelligence when a fund's process demands it. Until those hit, the free stack covers sourcing timing, which is where returns are made. The PitchBook alternatives and Crunchbase alternatives comparisons price the upgrade paths.
Key takeaways#
Sourcing timing is free; enrichment is paid. Build the stack in layers, coverage, verification, monitoring, tooling, and pay only when a job breaks. The two-hour weekly workflow on free public data outperforms an expensive stack used passively, because the edge is cadence, not subscription tier.
The source catalogue at a glance#
The core free sources worth knowing cold. GitHub itself: org pages, commit history, contributor graphs, dependency manifests, the primary coverage layer for anything open-source adjacent. The public signals API: hundreds of tracked startup orgs, 15 sectors, weekly momentum, no key required, plus CSV and an MCP server for agent runtimes. SEC EDGAR: Form D filings, the authoritative record of US rounds, free, lagging announcements by days to weeks. Crunchbase free tier: round and company verification with rate limits. Job boards, LinkedIn company pages, and careers pages: the hiring layer, two to eight weeks of lead time. Product surfaces: changelogs, docs, status pages, release feeds. Community: founder activity on X and Hacker News, conference programs. The 47-source catalogue indexes all of these with coverage notes and lag times.
Data hygiene rules#
Free data needs discipline paid data pretends to not need. Date every capture: slope is the signal, and slopes require timestamps, so log weekly snapshots rather than relying on live views. Record the source with the fact: "funding per Form D" and "funding per TechCrunch" have different reliability, and conflating them poisons benchmarks. Never mix bases: a company can have 40 GitHub contributors and 4 engineers; the numbers measure different things, and the metrics guide defines each precisely. And deduplicate entities ruthlessly: free sources are full of renamed companies and stale profiles, and one wrong join can double-count a company into a trend.
The two cultures: free-stack vs paid-stack investors#
There is a real cultural split in the market. Paid-stack investors treat data as a subscription problem and get depth, standardization, and support. Free-stack investors treat data as a workflow problem and get timing, flexibility, and zero marginal cost. For sourcing specifically, timing wins: the free engineering layer leads the paid funding databases by weeks, which is exactly the window where sourcing happens. The mature position uses both, free for discovery, paid for enrichment, and the signal tools comparison maps which tools belong on which side.