What is a NeoLab?
A NeoLab is a research-first AI company founded by senior researchers who left a frontier lab such as OpenAI, Google DeepMind, Anthropic or Meta. It raises hundreds of millions of dollars, often at a billion-dollar valuation, before it has revenue or a product. Thinking Machines Lab, Safe Superintelligence and Periodic Labs are examples. The term spread in late 2025; we now track 102.
There's a new kind of startup in the valley: the Neolab. 9 of these 10 "neolabs" have $1B+ valuations at the seed. They're all likely <$10M in revenue and founded by ex-model lab AI researchers.Deedy Das, Menlo Ventures, 14 Nov 2025 (source)
A Neolab is a pre-revenue scale startup working on long-term AI breakthroughs, usually with a $1B+ valuation. There are now 63 of them!Deedy Das, May 2026 (source)
Research-first companies that own a domain-biased corpus, run an integrated RL and verification loop perhaps inside a high-stakes domain, and could exit into public markets or a strategic M&A wave that is already underway.NEA, "The AI Neolab Wild West", 17 Jun 2026 (source)
The broad definition is when a top researcher from an AI lab like OpenAI, Anthropic, xAI, or DeepMind spin out either with a novel research idea or direction they want to pursue independently. They're able to pull a small research team with them and large amounts of capital are needed from the get-go.Kevin Tu, Signal Processing, 2 Feb 2026 (source)
The term surfaced in The Information's reporting in November 2025 ("Investors Chase Neolabs to Outflank OpenAI, Anthropic"), was popularised by Deedy Das's list, and by mid-2026 had its own trackers, critics and compute-crunch stories. The Information counted $2.5 billion committed to five labs in a month; Inc. counted more than $10 billion by December 2025; Deedy's May 2026 list had 63 labs; neolabs.fyi lists 105.
Three exits from stealth
This site is the tracking half of an essay on gtmascode.dev, NeoLabs: Three Exits from Stealth, Three Different Buyers. Its argument: a lab exits obscurity for three different audiences, talent, capital and customers, and one announcement cannot supply the same proof to each. A famous founder buys the first meeting; sustained demand needs an artifact someone outside can use. The essay maps eight frontier labs and fourteen category labs by public positioning. The tables here track the same three facts for every lab: what it has released, who can use it, and what evidence exists that users stay. The Status column is that judgement, and it is deliberately coarse: stealth, research only, partnerships, API, open model, product, licensed, acquired, merged.
How this is built
A static site with no framework and no runtime. Three datasets, all committed so a build is reproducible:
- data/labs.json is hand-maintained on purpose. Every funding round carries the URL it was read from and the build refuses to run on one that does not. A valuation with no priced round is marked as an estimate and links to the tracker it came from.
- data/news.json is built from SERP candidates plus a hand-written seed. Items are deduplicated by canonical URL, tagged with the labs they name, classified as funding, launch, deal, product, talent or analysis by keyword rules, and dated from the date the search engine prints.
- data/market_map.json is derived: every investor named on a round, the labs it backs, how many rounds it led, and an investor-by-topic matrix.
Investor names are normalised through one alias table (a16z, Andreessen and Andreessen Horowitz are one investor; NVentures is Nvidia). Founder pedigree is the prior employer named in the announcement, normalised the same way, so Google Brain and DeepMind both read as Google DeepMind.
The sourcing agents
The site refreshes itself every three days. A scheduled run executes the scripts below, then a Claude agent does the judgement calls the scripts refuse to make, then the build is gated, tested and deployed. Nothing an unverified scrape produced reaches the labs table.
- News sourcing (
scripts/discover_news.py). A fixed list of sector queries and one query per tracked lab run through Bright Data's SERP API in Google News mode, so every result is an article with a publication date. Site-restricted queries run as web searches and take their date from the result or the URL. Transient upstream errors are retried with backoff, then once more after a cooldown, and the run's success rate is recorded. It never writes to the newsfeed directly. - Newsfeed build (
scripts/build_news.py). Results are deduplicated by URL and by headline, tagged with the labs they name, classified by kind, and dated from the URL path when it carries one. An undated result is a page, not news, and is dropped. A lab whose name is also an English word (Magic, Harmonic, Reflection) is only tagged when the story is about AI. - Stealth-lab detection (
scripts/discover_labs.py). Reads the launch-shaped results and pulls a company name from headlines shaped like X emerges from stealth or X raises $N, stripping descriptions such as "London-based AI lab". Known labs are dropped; the rest become leads, proposed once and then promoted or dismissed. - Jobs (
scripts/refresh_jobs.py). Every lab whose board is on Ashby, Greenhouse, Lever or Gem is re-read from that system's public API, with each role's posting date. Careers pages on a lab's own site are read through Bright Data's Web Unlocker. A board that fails keeps its previous roles with a dated note, so a flaky source never empties the page. - Verification (a scheduled Claude agent; the brief is in
agents/funding-verifier.md). Promotes or dismisses leads, confirms new rounds against a second source, records talks as talks, and updates a lab's status when it is acquired, licensed or ships. - The gate (
scripts/gate.py). The run does not deploy if the lab count falls by more than a tenth, the newsfeed shrinks, or fewer than half the searches succeed. That pattern means a source changed shape, not that the world changed.
Every figure on the site is computed at build time from these files: the counts in the masthead, the sums per topic and per origin on the mafias, the capital led on the market map. Nothing is typed into a page.
Sourcing rules
- Nothing is asserted without a link to where it came from.
- "In talks" and "reportedly seeking" figures are shown as talks and are never summed into totals or the valuation column.
- Round sizes and valuations are as reported by the linked source. Where sources disagree, the lab page carries the disagreement in the notes rather than a quiet average.
- Company-reported revenue (Skild's $100M ARR, Tinker's "few hundred million") is labelled as company-reported.
- Lab and person entries carry only public professional information the lab or the press published. Removal and correction requests are honoured without discussion: open an issue.
Known gaps
- Chinese labs are listed only where a tracker already lists them (DeepSeek, StepFun, Galbot). Deedy Das excludes them; neolabs.fyi includes some. Coverage is thin.
- Several watchlist labs have a tracker valuation and no disclosed round. They are marked est. everywhere.
- 48 of 102 labs have no sourced round yet. The verification agent's queue starts there.
Elsewhere
- Robotics Papers, the sister site: reading clubs, papers and the people presenting them. The robotics labs here link across.
- neolabs.fyi, cleverhack's tracker and GGV's read of Deedy's map, the other trackers this one was checked against.
- Download the datasets as JSON.