New startup ideas · AI and software · Agent infrastructure

startup idea

Rubricon

An exchange where domain experts sell held-out eval suites that enterprises run on demand.

Rubricon is a marketplace for agent evaluation: radiologists, tax accountants, claims adjusters and network engineers publish task sets with graders, and enterprises pay per run to test their agents against them.

4/5

venture judge

3

similar startups, last 2 years (14 all-time)

83%

of 4 nearest real companies still alive

yes

3 matching federal grants and programs

Direction supported by government programs and grants

Test it before you build it

$1,900 · 6 weeks · 20 prospects

Prove for under $2,000 in six weeks that US claims teams will pay real money to run their AI agents against an adjuster-written eval suite, and that working adjusters will write one for a revenue share.

1Focus group: who and where

VP of Claims, claims innovation lead, or head of claims automation at a US P&C carrier or TPA (500-5,000 employees) that is piloting an AI claims agent this quarter and has no independent way to test it before it touches live claims; secondarily, product leads at claims-automation vendors who need third-party scores to close enterprise deals. Supply side: licensed adjusters with 5+ years of desk experience.

where to find 20 · LinkedIn Sales Navigator filtered to titles VP Claims, Claims Innovation, Head of Claims Automation at P&C carriers and TPAs with keyword AI; InsurTech NY events; InsureTech Connect (ITC Vegas, October 2026) attendee outreach; adjusters via r/adjusters and LinkedIn title searches for claims adjuster.

2Sell first, build later

A private benchmark of your claims agent against a 60-task suite written by 5 working adjusters: 3 scored runs, per-task grading with adjuster rationale, and a written report your model risk or vendor review team can file, delivered within 21 days of payment. The answer key never leaves our sandbox.

the ask · $2,000 per pilot benchmark (3 runs), creditable against $500 per run afterward

a real yes · A real yes is a paid Stripe invoice or a purchase order with a dated agent endpoint handoff. Compliments, requests to see the task list for free, and unpaid 'run it and show us' asks are not a yes.

3Small experiments

The first one attacks the riskiest assumption; each ends with a number that says whether to run the next.

  1. 1. Adjuster task sprint

    $1,000 · 10 days

    Recruit 5 licensed adjusters through r/adjusters and LinkedIn DMs and pay each $200 to write 20 claims decision tasks with grading rubrics under an NDA. This tests whether the supply side exists at a price the marketplace can bear and whether they will sign a revenue-share for future paid runs.

    keep going if · 4 of 5 adjusters deliver usable tasks with rubrics inside 10 days and agree in writing to a 30% revenue share on future runs

  2. 2. Concierge benchmark on calls

    $300 · 14 days

    Assemble the tasks into a 60-task adjuster suite, run two frontier models against it by hand, and walk the scored report through 15 booked calls with claims leads. The suite never leaves your machine; the report is a PDF. Ask at the end of each call: do you want your own agent scored against this.

    keep going if · 8 of 15 calls end with the prospect asking to run their own agent or vendor's agent against the suite

  3. 3. Priced landing page

    $450 · 14 days

    One-page site: adjuster-written claims agent benchmark, $500 per evaluation run, first cohort of 5 buyers, Stripe Checkout behind the button. Drive 200 targeted visits with $300 of LinkedIn ads against the Sales Navigator audience and direct outreach.

    keep going if · 10 qualified signups with work emails and 3 who start Stripe Checkout

4Collect a deposit up front

Tesla took $1,000 refundable reservations for the Model 3 and $100 for the Cybertruck before building either: the deposit is the measurement, not the revenue.

$2,000

per prospect, refundable

how · Paid design-partner pilot invoiced 100% up front through a Stripe invoice, with a one-page order form signed by the claims or vendor-management lead set up: Stripe Invoicing

what it reserves · One of 5 first-cohort slots, per-run pricing locked at $500 for 12 months, and a vote on the next 40 tasks added to the suite

refund · Full refund if the scored report is not delivered within 21 days of the agent endpoint being connected

target · 3 paid pilots ($6,000 collected) from 20 conversations within 5 weeks

Go: build it if

3 pilots paid up front, 8 of 15 calls request a scored run, and 4 of 5 adjusters signed to a revenue share: build the sandbox.

Kill: stop if

Fewer than 1 paid pilot after 20 completed conversations, or fewer than 5 of 15 calls ask for a run: claims teams are content with vendor-supplied evals, stop.

Would you run this test?

One tap. The yes-share feeds the Demand pillar of this idea's score; nobody sees who answered.

Budgets are out-of-pocket estimates for a team of one to three, US market. Size the deposit to the deal, and check the terms before taking money in a regulated line.

Scorecard

One score that balances how trendy the idea is, the demand for it and its potential for 100x, with competition measured relative to every other idea in the catalog. Recent startup trends first, government priorities second.

76

Idea Score, 0-100 · raw 46.0 x 1.64

Warm

competition: more crowded than 41% of ideas · headwind x0.80

+1.5

government priorities, secondary (3 matching grants)

Trend

51

Is the wave forming now? 2025-26 entrants vs 2023-24, rounds since 2025, the sector's live-batch direction, the 2026 trend analyst.

  • Entrants 2025-26 vs 2023-24 (similar companies)70
  • Rounds announced 2025+ in the sector8
  • Sector direction (live batch)100
  • 2026 trend analyst25

Demand

52

Does anyone want it? YC's current RFS, companies already paid for something similar, the operator judge, founders' yes-rate in decks, readers who would run the test.

  • YC asks for it (current RFS: idea / sector)30
  • Someone already pays (similar companies, recent / all-time)100
  • Operator judge: real pain25

100x potential

67

Can it return a fund? The venture judge (double weight), market-size and moat axes, neighbours still alive, the technologist judge.

  • Venture judge75
  • Market size axis67
  • Moat axis100
  • Neighbours still alive34
  • Technologist judge50

Score = 100 x cbrt(Trend x Demand x 100x) x (1 - 0.5 x crowding) + government bonus (max 5), calibrated so the 95th-percentile idea scores 90 (order never changes). A geometric mean: a weak pillar cannot be papered over. Percentiles are among the 272 ideas in the catalog; the terms matched were exchange, domain, experts, sell, held-out, eval, suites, marketplace.

The idea in full

What
Rubricon is a marketplace for agent evaluation: radiologists, tax accountants, claims adjusters and network engineers publish task sets with graders, and enterprises pay per run to test their agents against them. Sellers never release the answer keys, since suites run only inside Rubricon's sandbox with private held-out splits, and sellers take a revenue share on every run. Buyers sign up with a card and point their agent endpoint at a suite in an afternoon.
Why now
CoArena is building the biggest crowdsourced benchmark for computer-use, and the NSF funded an I-Corps translation project for a GenAI result comparison tool for evaluating foundation model outputs at roughly $50k, so third-party comparison of agent outputs is becoming a traded good; meanwhile Atla and Baserun both died selling single-vendor eval platforms.
Wedge: first customer and entry point
US insurance claims teams buying one suite: adjuster decision tasks written by ten working adjusters, priced per run, sold self-serve.
Path to 100x
Enterprise assurance and QA spend moving onto agent systems is a $10-100B pool, and an exchange wins most of it because the suite with the most buyers attracts the best experts, whose suites then attract more buyers. Once procurement teams cite Rubricon scores in agent purchases, the scoreboard itself is the category and a second exchange has nothing to compare against.
Ceiling
Enterprises keep accepting vendor-supplied evals and internal spreadsheets, leaving Rubricon a niche developer tool at low tens of millions in revenue.
Closest real companies, as the generator saw them
CoArena crowdsources one horizontal computer-use benchmark; Rubricon is a two-sided market of priced vertical suites. Buildbox and Agnost AI analyze how users experience one company's agent rather than comparing agents across a shared standard. Atla and Baserun were single-vendor eval tooling, not an exchange.
Main risk
Suites leak into training data or public repositories and lose their value faster than sellers can replace them.

Five judges

Each judge scores every idea in the catalog with a named rubric; the venture judge decides whether a card is shown at all (4-5 is venture-grade).

  • Venture investor

    4/5

    Capital-light exchange with revenue within a year where the scoreboard cited in procurement compounds sellers and buyers into one defensible standard.

  • Bootstrapper

    4/5

    Ten working adjusters writing one suite, sold self-serve per run within a year, is cheap to launch and cash-positive fast.

  • Operator

    2/5

    Atla and Baserun both died proving nobody has budget for eval, and procurement citing Rubricon scores requires the market to move first.

  • Technologist

    3/5

    Held-out sandbox execution is modest engineering and suite leakage into training data erodes the asset faster than the network effect builds.

  • Risk

    3/5

    Capital light with no platform dependency, but suite value evaporates the moment held-out splits leak into training data.

  • trends

    2/5

    Paid held-out eval suites cite only a $50k NSF award, and the graveyard entries Atla and Baserun show the eval trend already tried and thinned.

Similar startups in the directory

Companies whose pitch matches most of the idea's terms (exchange, domain, experts, sell, held-out, eval, suites, marketplace): 14 all-time, 3 from the last two years. Same matching as Idea Check.

  • Robocurveyc S26 · 2026 · Robotics and physical worldalive

    Evaluating robots in the real world.

  • Proxyc F25 · 2025 · Vertical AI agentsalive

    AI technical support for complex physical products

  • Misprintyc W25 · 2025 · Consumeralive

    Misprint is building Robinhood for Pokemon cards

  • MotorMiaseedcamp Seedcamp 2023 · 2023 · Horizontal AI assistantsalive

    Marketplace for automotive enthusiasts to buy and sell speciality equipment

  • Mageyc W19 · 2019 · Consumersite down

    The safest way to buy and sell collectibles

  • Vertoyc W19 · 2019 · Fintechalive

    Cross-Border Payments and FX Simplified

  • Dabchy500global · 2018 · Consumeralive

    Operator of a peer-to-peer fashion marketplace intended to connect buyers and sellers for online shopping. The company offers a marketplace that is a mix of a social network and a forum where people like and comment on the articles posted and give each other fashion tips, enabling customers to buy, sell, exchange, and review new and used fashion items with ease.

  • allritessosv SOSV Orbit Startups - Chinaccelerator 15 · 2017 · Commerce and marketplacesalive

    Asia’s leading media marketplace in the US and SEA to Buy & Sell Film, TV & Sports Content

  • Ottimate500global · 2015 · Vertical AI agentsalive

    Operator of artificial intelligence (AI)-powered marketplace intended to connect finance and accounting freelance professionals with companies. The company's platform matches businesses of all sizes and industries with highly vetted finance experts, Chief Financial Officers (CFOs), Certified Public Accountants (CPAs), and bookkeepers who have the domain expertise to tackle industry and company-specific problems on demand, enabling businesses to access the tools and insights they need to manage and grow their books of business.

  • SupplyHog Holdings, LLC.500global 500G GA 5 · 2012 · Commerce and marketplacesacquired

    Provider of an online marketplace designed to sell construction supplies. The company's online marketplace offers drop shipping and contractor-level pricing services along with hardware and building supplies, enabling general contractors to procure building materials.

  • Cardpoolyc W10 · 2010 · Consumeracquired
  • GearUpplugandplay · Commerce and marketplacesunchecked

    GearUp is a peer to peer marketplace for creatives to buy and sell new and used creative equipment.

Run this as an Idea Check →

The generator's reference companies

Real companies the model named as closest when it wrote the card, with their fate. A check mark is a company the radar could verify in its directory.

Public money in this direction

US federal grants, SBIR/STTR awards and open opportunities from the radar's public-money feed, matched to the idea's terms; the sector totals give the context.

3

grants and programs matching the idea

35

startup-relevant grants in Agent infrastructure

$15M

awarded in the sector, tracked

1

opportunities open now in the sector

All public money by sector →

Market signal

What the radar sees in Agent infrastructure: new companies by cohort year, the forming YC batch, and outcomes since the February snapshot.

Agent infrastructure · 27 → 63 → 106 → 148 → 126 new companies 2022 → 2026 · 95% aliveYC S26: 39 in this cluster, 16% of the batch (was 17% in X26) (F26 is still forming: 21 listed)Since February, of 244 YC companies here: 7 acquired, 2 shut down, 102 rewrote their pitch

Agent infrastructure: companies, trend and grants →

Design attributes

The card is one cell of a designed set: every axis below was chosen before the text was written, and the text had to realize it.

Buyer
Enterprise
Business model
Marketplace
Path to 100x
Creates a new category
Market size
$10-100B market
Capital intensity
Capital-light (software margins)
Speed to revenue
Revenue within a year
Technical depth
Real engineering
Go-to-market
Self-serve
Moat
Network effects
Geography
US first
Regulation
Some regulation
Vibe
Boring business

Listed under

An idea sits in its own sector and in any sector its text clearly touches.

More ideas like this

AI and software · Agent infrastructure

Socketry

The exchange where software vendors sell maintained, agent-ready API connections.

Socketry is a two-sided marketplace where SaaS vendors publish guaranteed-current, agent-callable versions of their APIs - tested sandboxes, auth, rate contracts, change notices - and enterprises subscribe to them for their internal agents with one bill and one security review.

Score 88Open competitionVC 4/5MarketplaceEnterprise98% of 4 neighbours alive

AI and software · Agent infrastructure

Latchwork

Self-healing connectors for the long tail of small business software agents cannot reach.

Latchwork records a session against a vertical tool that has no usable API - a dental scheduler, a salon booking system, a freight dispatch app - and turns it into a connector that agents call like an API, then repairs itself when the vendor changes a screen.

Score 88Open competitionVC 4/5Software subscriptionSmall businesstest: $1.6k · 6w98% of 4 neighbours alive

AI and software · Agent infrastructure

Civitrace

Casework agents for US agencies that turn every verified fact into reusable public evidence.

Civitrace runs eligibility and permit casework as a managed agent service for US state, county and city agencies: it reads submitted documents, checks them against source systems, drafts determinations, and hands a human caseworker a decision packet with citations.

Score 75Open competitionVC 3/5AI agent as a serviceGovernment and public sector98% of 4 neighbours alive

AI and software · Agent infrastructure

Actledger

Signed action records for enterprise agents, with a policy pack ecosystem on top.

Actledger sits between an enterprise's agents and the systems they touch: every tool call is authorized against policy, signed, and written to an immutable action record mapped to audit controls.

Score 69Active competitionVC 4/5Software subscriptionEnterprisetest: $1.4k · 6w98% of 4 neighbours alive

AI and software · Agent infrastructure

Toolharbor

The governed registry every enterprise agent must pass through to touch a tool.

Toolharbor is a control plane that sits between an enterprise's AI agents and every tool, API, and MCP server they call, enforcing policy, credentials, and rate limits per agent.

Score 68Active competitionVC 5/5Software subscriptionEnterprise97% of 3 neighbours alive

AI and software · Agent infrastructure

Openstall

No-code layer that makes every small business readable and transactable for AI agents.

An SMB connects its booking, inventory, and payment tools (Shopify, Square, Calendly, QuickBooks) in a no-code dashboard, and Openstall publishes them as one hosted, self-maintaining endpoint that any AI agent can query and transact against.

Score 66Active competitionVC 4/5Software subscriptionSmall business97% of 3 neighbours alive
Swipe ideas like this in the deckTalk to the radar about it

Fictional company written 2026-08-22 from MarkosWeb data; the companies, grants and numbers around it are real and tracked. Treat the idea as a research prompt, not a plan.