New startup ideas · AI and software · Agent infrastructure
startup idea
Rubricon
An exchange where domain experts sell held-out eval suites that enterprises run on demand.
Rubricon is a marketplace for agent evaluation: radiologists, tax accountants, claims adjusters and network engineers publish task sets with graders, and enterprises pay per run to test their agents against them.
- Marketplace
- Enterprise
- $10-100B market
- Creates a new category
- US first
4/5
venture judge
3
similar startups, last 2 years (14 all-time)
83%
of 4 nearest real companies still alive
yes
3 matching federal grants and programs
Direction supported by government programs and grants
Test it before you build it
$1,900 · 6 weeks · 20 prospects
Prove for under $2,000 in six weeks that US claims teams will pay real money to run their AI agents against an adjuster-written eval suite, and that working adjusters will write one for a revenue share.
1Focus group: who and where
VP of Claims, claims innovation lead, or head of claims automation at a US P&C carrier or TPA (500-5,000 employees) that is piloting an AI claims agent this quarter and has no independent way to test it before it touches live claims; secondarily, product leads at claims-automation vendors who need third-party scores to close enterprise deals. Supply side: licensed adjusters with 5+ years of desk experience.
where to find 20 · LinkedIn Sales Navigator filtered to titles VP Claims, Claims Innovation, Head of Claims Automation at P&C carriers and TPAs with keyword AI; InsurTech NY events; InsureTech Connect (ITC Vegas, October 2026) attendee outreach; adjusters via r/adjusters and LinkedIn title searches for claims adjuster.
2Sell first, build later
A private benchmark of your claims agent against a 60-task suite written by 5 working adjusters: 3 scored runs, per-task grading with adjuster rationale, and a written report your model risk or vendor review team can file, delivered within 21 days of payment. The answer key never leaves our sandbox.
the ask · $2,000 per pilot benchmark (3 runs), creditable against $500 per run afterward
a real yes · A real yes is a paid Stripe invoice or a purchase order with a dated agent endpoint handoff. Compliments, requests to see the task list for free, and unpaid 'run it and show us' asks are not a yes.
3Small experiments
The first one attacks the riskiest assumption; each ends with a number that says whether to run the next.
1. Adjuster task sprint
$1,000 · 10 days
Recruit 5 licensed adjusters through r/adjusters and LinkedIn DMs and pay each $200 to write 20 claims decision tasks with grading rubrics under an NDA. This tests whether the supply side exists at a price the marketplace can bear and whether they will sign a revenue-share for future paid runs.
keep going if · 4 of 5 adjusters deliver usable tasks with rubrics inside 10 days and agree in writing to a 30% revenue share on future runs
2. Concierge benchmark on calls
$300 · 14 days
Assemble the tasks into a 60-task adjuster suite, run two frontier models against it by hand, and walk the scored report through 15 booked calls with claims leads. The suite never leaves your machine; the report is a PDF. Ask at the end of each call: do you want your own agent scored against this.
keep going if · 8 of 15 calls end with the prospect asking to run their own agent or vendor's agent against the suite
3. Priced landing page
$450 · 14 days
One-page site: adjuster-written claims agent benchmark, $500 per evaluation run, first cohort of 5 buyers, Stripe Checkout behind the button. Drive 200 targeted visits with $300 of LinkedIn ads against the Sales Navigator audience and direct outreach.
keep going if · 10 qualified signups with work emails and 3 who start Stripe Checkout
4Collect a deposit up front
Tesla took $1,000 refundable reservations for the Model 3 and $100 for the Cybertruck before building either: the deposit is the measurement, not the revenue.
$2,000
per prospect, refundable
how · Paid design-partner pilot invoiced 100% up front through a Stripe invoice, with a one-page order form signed by the claims or vendor-management lead set up: Stripe Invoicing ↗
what it reserves · One of 5 first-cohort slots, per-run pricing locked at $500 for 12 months, and a vote on the next 40 tasks added to the suite
refund · Full refund if the scored report is not delivered within 21 days of the agent endpoint being connected
target · 3 paid pilots ($6,000 collected) from 20 conversations within 5 weeks
Go: build it if
3 pilots paid up front, 8 of 15 calls request a scored run, and 4 of 5 adjusters signed to a revenue share: build the sandbox.
Kill: stop if
Fewer than 1 paid pilot after 20 completed conversations, or fewer than 5 of 15 calls ask for a run: claims teams are content with vendor-supplied evals, stop.
Would you run this test?
One tap. The yes-share feeds the Demand pillar of this idea's score; nobody sees who answered.
Budgets are out-of-pocket estimates for a team of one to three, US market. Size the deposit to the deal, and check the terms before taking money in a regulated line.
Scorecard
One score that balances how trendy the idea is, the demand for it and its potential for 100x, with competition measured relative to every other idea in the catalog. Recent startup trends first, government priorities second.
76
Idea Score, 0-100 · raw 46.0 x 1.64
Warm
competition: more crowded than 41% of ideas · headwind x0.80
+1.5
government priorities, secondary (3 matching grants)
Trend
51
Is the wave forming now? 2025-26 entrants vs 2023-24, rounds since 2025, the sector's live-batch direction, the 2026 trend analyst.
- Entrants 2025-26 vs 2023-24 (similar companies)70
- Rounds announced 2025+ in the sector8
- Sector direction (live batch)100
- 2026 trend analyst25
Demand
52
Does anyone want it? YC's current RFS, companies already paid for something similar, the operator judge, founders' yes-rate in decks, readers who would run the test.
- YC asks for it (current RFS: idea / sector)30
- Someone already pays (similar companies, recent / all-time)100
- Operator judge: real pain25
100x potential
67
Can it return a fund? The venture judge (double weight), market-size and moat axes, neighbours still alive, the technologist judge.
- Venture judge75
- Market size axis67
- Moat axis100
- Neighbours still alive34
- Technologist judge50
Score = 100 x cbrt(Trend x Demand x 100x) x (1 - 0.5 x crowding) + government bonus (max 5), calibrated so the 95th-percentile idea scores 90 (order never changes). A geometric mean: a weak pillar cannot be papered over. Percentiles are among the 272 ideas in the catalog; the terms matched were exchange, domain, experts, sell, held-out, eval, suites, marketplace.
The idea in full
- What
- Rubricon is a marketplace for agent evaluation: radiologists, tax accountants, claims adjusters and network engineers publish task sets with graders, and enterprises pay per run to test their agents against them. Sellers never release the answer keys, since suites run only inside Rubricon's sandbox with private held-out splits, and sellers take a revenue share on every run. Buyers sign up with a card and point their agent endpoint at a suite in an afternoon.
- Why now
- CoArena is building the biggest crowdsourced benchmark for computer-use, and the NSF funded an I-Corps translation project for a GenAI result comparison tool for evaluating foundation model outputs at roughly $50k, so third-party comparison of agent outputs is becoming a traded good; meanwhile Atla and Baserun both died selling single-vendor eval platforms.
- Wedge: first customer and entry point
- US insurance claims teams buying one suite: adjuster decision tasks written by ten working adjusters, priced per run, sold self-serve.
- Path to 100x
- Enterprise assurance and QA spend moving onto agent systems is a $10-100B pool, and an exchange wins most of it because the suite with the most buyers attracts the best experts, whose suites then attract more buyers. Once procurement teams cite Rubricon scores in agent purchases, the scoreboard itself is the category and a second exchange has nothing to compare against.
- Ceiling
- Enterprises keep accepting vendor-supplied evals and internal spreadsheets, leaving Rubricon a niche developer tool at low tens of millions in revenue.
- Closest real companies, as the generator saw them
- CoArena crowdsources one horizontal computer-use benchmark; Rubricon is a two-sided market of priced vertical suites. Buildbox and Agnost AI analyze how users experience one company's agent rather than comparing agents across a shared standard. Atla and Baserun were single-vendor eval tooling, not an exchange.
- Main risk
- Suites leak into training data or public repositories and lose their value faster than sellers can replace them.
Five judges
Each judge scores every idea in the catalog with a named rubric; the venture judge decides whether a card is shown at all (4-5 is venture-grade).
Venture investor
4/5
Capital-light exchange with revenue within a year where the scoreboard cited in procurement compounds sellers and buyers into one defensible standard.
Bootstrapper
4/5
Ten working adjusters writing one suite, sold self-serve per run within a year, is cheap to launch and cash-positive fast.
Operator
2/5
Atla and Baserun both died proving nobody has budget for eval, and procurement citing Rubricon scores requires the market to move first.
Technologist
3/5
Held-out sandbox execution is modest engineering and suite leakage into training data erodes the asset faster than the network effect builds.
Risk
3/5
Capital light with no platform dependency, but suite value evaporates the moment held-out splits leak into training data.
trends
2/5
Paid held-out eval suites cite only a $50k NSF award, and the graveyard entries Atla and Baserun show the eval trend already tried and thinned.
Similar startups in the directory
Companies whose pitch matches most of the idea's terms (exchange, domain, experts, sell, held-out, eval, suites, marketplace): 14 all-time, 3 from the last two years. Same matching as Idea Check.
Evaluating robots in the real world.
AI technical support for complex physical products
Misprint is building Robinhood for Pokemon cards
Marketplace for automotive enthusiasts to buy and sell speciality equipment
The safest way to buy and sell collectibles
Cross-Border Payments and FX Simplified
Operator of a peer-to-peer fashion marketplace intended to connect buyers and sellers for online shopping. The company offers a marketplace that is a mix of a social network and a forum where people like and comment on the articles posted and give each other fashion tips, enabling customers to buy, sell, exchange, and review new and used fashion items with ease.
Asia’s leading media marketplace in the US and SEA to Buy & Sell Film, TV & Sports Content
Operator of artificial intelligence (AI)-powered marketplace intended to connect finance and accounting freelance professionals with companies. The company's platform matches businesses of all sizes and industries with highly vetted finance experts, Chief Financial Officers (CFOs), Certified Public Accountants (CPAs), and bookkeepers who have the domain expertise to tackle industry and company-specific problems on demand, enabling businesses to access the tools and insights they need to manage and grow their books of business.
Provider of an online marketplace designed to sell construction supplies. The company's online marketplace offers drop shipping and contractor-level pricing services along with hardware and building supplies, enabling general contractors to procure building materials.
GearUp is a peer to peer marketplace for creatives to buy and sell new and used creative equipment.
The generator's reference companies
Real companies the model named as closest when it wrote the card, with their fate. A check mark is a company the radar could verify in its directory.
Public money in this direction
US federal grants, SBIR/STTR awards and open opportunities from the radar's public-money feed, matched to the idea's terms; the sector totals give the context.
3
grants and programs matching the idea
35
startup-relevant grants in Agent infrastructure
$15M
awarded in the sector, tracked
1
opportunities open now in the sector
- Collaborative Research: Cognitively Optimized Mathematics through Peer Analogous Reflective Exchangesawardmedium relevance
National Science Foundation · Cyberlearn & Future Learn Tech · $506K · posted 2026-08-12
- Collaborative Research: Cognitively Optimized Mathematics through Peer Analogous Reflective Exchangesawardmedium relevance
National Science Foundation · Cyberlearn & Future Learn Tech · $224K · posted 2026-08-12
National Science Foundation · Secure &Trustworthy Cyberspace · $200K · posted 2026-08-07
Market signal
What the radar sees in Agent infrastructure: new companies by cohort year, the forming YC batch, and outcomes since the February snapshot.
Agent infrastructure · 27 → 63 → 106 → 148 → 126 new companies 2022 → 2026 · 95% aliveYC S26: 39 in this cluster, 16% of the batch (was 17% in X26) (F26 is still forming: 21 listed)Since February, of 244 YC companies here: 7 acquired, 2 shut down, 102 rewrote their pitch
Design attributes
The card is one cell of a designed set: every axis below was chosen before the text was written, and the text had to realize it.
- Buyer
- Enterprise
- Business model
- Marketplace
- Path to 100x
- Creates a new category
- Market size
- $10-100B market
- Capital intensity
- Capital-light (software margins)
- Speed to revenue
- Revenue within a year
- Technical depth
- Real engineering
- Go-to-market
- Self-serve
- Moat
- Network effects
- Geography
- US first
- Regulation
- Some regulation
- Vibe
- Boring business
Listed under
An idea sits in its own sector and in any sector its text clearly touches.
More ideas like this
AI and software · Agent infrastructure
Socketry
The exchange where software vendors sell maintained, agent-ready API connections.
Socketry is a two-sided marketplace where SaaS vendors publish guaranteed-current, agent-callable versions of their APIs - tested sandboxes, auth, rate contracts, change notices - and enterprises subscribe to them for their internal agents with one bill and one security review.
AI and software · Agent infrastructure
Latchwork
Self-healing connectors for the long tail of small business software agents cannot reach.
Latchwork records a session against a vertical tool that has no usable API - a dental scheduler, a salon booking system, a freight dispatch app - and turns it into a connector that agents call like an API, then repairs itself when the vendor changes a screen.
AI and software · Agent infrastructure
Civitrace
Casework agents for US agencies that turn every verified fact into reusable public evidence.
Civitrace runs eligibility and permit casework as a managed agent service for US state, county and city agencies: it reads submitted documents, checks them against source systems, drafts determinations, and hands a human caseworker a decision packet with citations.
AI and software · Agent infrastructure
Actledger
Signed action records for enterprise agents, with a policy pack ecosystem on top.
Actledger sits between an enterprise's agents and the systems they touch: every tool call is authorized against policy, signed, and written to an immutable action record mapped to audit controls.
AI and software · Agent infrastructure
Toolharbor
The governed registry every enterprise agent must pass through to touch a tool.
Toolharbor is a control plane that sits between an enterprise's AI agents and every tool, API, and MCP server they call, enforcing policy, credentials, and rate limits per agent.
AI and software · Agent infrastructure
Openstall
No-code layer that makes every small business readable and transactable for AI agents.
An SMB connects its booking, inventory, and payment tools (Shopify, Square, Calendly, QuickBooks) in a no-code dashboard, and Openstall publishes them as one hosted, self-maintaining endpoint that any AI agent can query and transact against.
Fictional company written 2026-08-22 from MarkosWeb data; the companies, grants and numbers around it are real and tracked. Treat the idea as a research prompt, not a plan.