New startup ideas · AI for people who run the AI themselves · Evaluation and trust the user controls

startup concept

Skillbench

Regression testing for Agent Skills before every model and skill update

A hosted test harness for Agent Skills.

54

similar startups, last 2 years (143 all-time)

yes

8 matching federal grants and programs

Direction supported by government programs and grants

Test it before you build it

$270 · 3 weeks · 25 prospects

For $270 and 3 weeks, prove that Agent Skills authors will pay cash today for regression tests of their SKILL.md, starting with a $49 manual run.

Riskiest assumption · Individual skill authors who distribute skills to others exist in paying numbers and will pay $30 a month for third-party regression testing, rather than eyeballing a few prompts after each update and waiting for agentskills.io or a major runtime to ship a free conformance suite.

1Focus group: who and where

An individual author who publishes or sells Agent Skills - a public repo with a SKILL.md and real users - pays $100 to $200 a month for a Max-tier seat, and watched a skill misbehave after a model or runtime update in the past 90 days.

where to find 25 · GitHub code search for repos containing SKILL.md sorted by stars and recent activity (a live directory of exactly these authors), the official Claude Developers Discord, skill-sharing threads on r/ClaudeAI, and author pages on agentskills.io.

2Sell first, build later

Today: a $49 manual regression run - ten fixture tasks built from your transcripts, replayed against Sonnet 5 and Haiku 4.5, line-by-line diff report in 48 hours. By October 15, 2026: the hosted harness that reruns your pinned suite automatically on every model version and skill edit.

the ask · $49 per manual run now; $99 founding year for the hosted harness (then $30 a month)

a real yes · A real yes is $49 or $99 actually charged. 'I'd try the free tier', GitHub stars, and offers to beta test for free do not count.

3Small experiments

The first one attacks the riskiest assumption; each ends with a number that says whether to run the next.

  1. 1. Fifteen author regression interviews

    $0 · 10 days

    Pull the 40 most-starred, recently updated SKILL.md repos from GitHub code search and contact the authors through repo issues, Discord, or their listed contact. Book 15 twenty-minute calls, get the story of their last model-update breakage in specifics, and close each call with a money ask: $49 for a manual regression run of their skill, ten fixtures, two models, results in 48 hours.

    keep going if · 9 of 15 shipped a regression after a model update in the past 90 days, and 5 of 15 pay the $49 on the call

  2. 2. Paid manual regression runs

    $150 · 12 days

    Deliver every purchased run by hand: write ten fixture tasks from the author's own transcripts and examples, replay the skill against Sonnet 5 and Haiku 4.5 via the API (about $150 of model spend across runs), and send a line-by-line diff report. Manual delivery is the right first artifact here because the product's entire value is this report, and producing it with the models directly is faster than any prototype harness.

    keep going if · 5 of 20 authors asked buy the $49 run, and 3 of 5 buyers say they would rerun it before every model update

  3. 3. Founding-author pre-order page

    $120 · 14 days

    Put up a one-page site with a redacted diff report from a real run, the $99 founding year (regular $360 at $30 a month), delivery October 15, 2026, and card checkout. Send it to every interviewee and buyer, and post it once in the Claude Developers Discord and once on r/ClaudeAI.

    keep going if · 8 of 25 authors asked pay the $99 founding year

4Collect a deposit up front

Tesla took $1,000 refundable reservations for the Model 3 and $100 for the Cybertruck before building either: the deposit is the measurement, not the revenue.

$99

per prospect, refundable

how · Two rungs of real money by card checkout: the $49 manual run (money moved today for the exact artifact the product automates) and the $99 refundable founding-year pre-order. Card checkout fits a solo author buying a sub-$100 developer tool online with nobody else to sign off. set up: Stripe Checkout

what it reserves · The founding price of $99 a year for life of the subscription (versus $360), a slot in the first cohort on October 15, 2026, and their fixture suite migrated into the harness for free

refund · The $99 is fully refundable by email until launch on October 15, 2026 and for 30 days after; delivered $49 runs are not refundable.

target · 5 paid $49 runs and 8 paid $99 pre-orders from 25 authors within 21 days

Go: build it if

5 or more $49 runs sold, 8 or more $99 pre-orders, and 3 of 5 run-buyers saying they would rerun before every model update: build the harness.

Kill: stop if

Fewer than 3 paid runs after 25 asks, or fewer than 5 of 15 interviewed authors report any regression in the past 90 days: stop, the pain is too rare to beat a free native conformance suite when one ships.

5 Scripts to run itoutreach message, landing copy, deposit terms · click to open

outreach message

Your skill has real users on GitHub, which means every model update can silently break it for all of them before you notice. I'm building Skillbench: point it at your SKILL.md, record ten fixture tasks, and it replays the suite against every new model version, flagging drift line by line. Before writing the product I'm doing paid manual runs: $49, ten fixtures from your own transcripts, two models, diff report back in 48 hours. Got 20 minutes this week to scope yours?

landing page

Know your skill still works before your users find out $49 per manual regression run today; $99 founding year for the hosted harness (then $360) Buy a run or reserve your founding year - $99, refundable until launch on October 15, 2026

deposit terms

Your $99 buys the founding year of Skillbench, live October 15, 2026, and locks your price at $99 a year instead of $360; we also migrate your fixture suite in for free. Fully refundable by email until launch and for 30 days after. Manual $49 runs are delivered within 48 hours and are not refundable once delivered.

Would you run this test?

One tap. The yes-share feeds the Demand pillar of this idea's score; nobody sees who answered.

Budgets are out-of-pocket estimates for a team of one to three, US market. Size the deposit to the deal, and check the terms before taking money in a regulated line.

Scorecard

Ranked against every idea in the catalog: trend, demand and 100x potential from the corpus, competition relative to the other ideas. A generated concept has no judges or swipes yet, so its pillars use the data signals only.

59

Idea Score, 0-100 · raw 36.5 x 1.61

Crowded

competition: more crowded than 98% of ideas · headwind x0.51

+5.0

government priorities, secondary (667 matching grants)

Trend

55

Is the wave forming now? 2025-26 entrants vs 2023-24, rounds since 2025, the sector's live-batch direction, the 2026 trend analyst.

  • Entrants 2025-26 vs 2023-24 (similar companies)70
  • Rounds announced 2025+ in the sector0
  • Rounds announced 2025+ matching the idea99
  • Sector direction (live batch)50

Demand

65

Does anyone want it? YC's current RFS, companies already paid for something similar, the operator judge, founders' yes-rate in decks, readers who would run the test.

  • YC asks for it (current RFS: idea / sector)30
  • Someone already pays (similar companies, recent / all-time)100

100x potential

67

Can it return a fund? The venture judge (double weight), market-size and moat axes, neighbours still alive, the technologist judge.

  • Neighbours still alive67

Score = 100 x cbrt(Trend x Demand x 100x) x (1 - 0.5 x crowding) + government bonus (max 5), calibrated so the 95th-percentile idea scores 90 (order never changes). A geometric mean: a weak pillar cannot be papered over. Percentiles are among the 382 ideas in the catalog; the terms matched were regression, testing, skills, model, skill, hosted, test, harness.

The concept in full

What
A hosted test harness for Agent Skills. A builder points it at a SKILL.md folder, records fixture tasks with expected outputs, and Skillbench replays the suite against new model versions and skill edits, flagging behavior drift line by line. In the first hour a user imports one skill, auto-generates ten fixture cases from past transcripts, and gets a pass or fail report they can pin to a version.
Grounded in (2025-2026 signals)
'Agent Skills package procedural knowledge as folders with a SKILL.md' launched October 16, 2025, 'released as an open standard on December 18, 2025 (agentskills.io), and by June 2026 about 40 products supported it, including Claude, OpenAI Codex, GitHub Copilot, VS Code, Cursor, Gemini CLI and Goose'. Claude Max launched April 9, 2025 at $100 and $200 a month for people who run the model all day.
What it rides
Agent Skills: launched October 16, 2025, an open standard since December 18. Skills are now versioned artifacts running across many assistants, so a change in the skill or the underlying model silently changes results everywhere at once, and nothing in the standard tests for that.
Why now
About 40 products supported the Agent Skills standard by June 2026, eight months after the October 16, 2025 launch, so one skill now runs on many runtimes and models with no shared test surface; the people maintaining those skills already pay $100 to $200 a month for Max-tier seats and will pay $30 more to stop shipping regressions.
Wedge: first customer and entry point
Individual skill authors selling or sharing skills built with the skill-creator skill; entry point is a free tier that tests one skill against two models, converting authors whose skill broke on a model update last month.
Closest real companies, as the generator saw them
Confident AI (yc W25) is an LLM eval and observability platform aimed at engineering teams instrumenting their own apps; Skillbench tests the SKILL.md artifact itself across third-party runtimes for a single author. Quantstruct (yc W25) tests stale product docs, not executable skills.
Main risk
The agentskills.io standard body or a major runtime ships a built-in conformance test suite and authors never look for a third-party one.

Similar startups in the directory

Companies whose pitch matches most of the concept's terms (regression, testing, skills, model, skill, hosted, test, harness).

  • アイリス株式会社plugandplay · Healthcare and bioalive

    Aillis specializes in the development of testing methods for Influenza using artificial intelligence technology.

  • Mirrorsyc F26 · 2026 · Agent infrastructurealive

    Catch and fix AI agent regressions before they reach production

  • Shepherd Roboticsyc F26 · 2026 · Robotics and physical worldalive

    Robots for high skilled labor powering AI infrastructure

  • BentoLabs AIyc X26 · 2026 · Agent infrastructurealive

    Monitoring and learning layer for long-running agents

  • Modayc W26 · 2026 · Agent infrastructurealive

    Harness Engineering Infrastructure that developers love.

  • Jobo Interimplugandplay PnP 2024 · 2024 · Commerce and marketplacesalive

    Jobo is a platform in Ivory Cost offering rapid, reliable, and automated temporary employment services. They use AI to preselect candidates within 24 hours and ensure all temporary workers are trained before starting their roles.

  • Haplotype Labsyc W24 · 2024 · Healthcare and bioalive

    Personalized prevention using genetics and AI

  • Kings500global 500G Georgia 3 · 2023 · Consumeralive

    Operator of an online learning platform intended to improve students' academic skills. The company provides an online diagnostic test to analyze student performance and provide information about their academic abilities as well as collaboration tools to allow students to work together on projects and assignments, enabling students to learn at their own pace with the opportunity to interact with other students and instructors to get feedback on their work.

  • OneSchemayc S21 · 2021 · Vertical AI agentsalive

    The AI Cowork for big data

  • CSPAyc S18 · 2018 · Developer toolsacquired

    The Computer Science Proficiency Assessment (CSPA™) is a…

  • Adafaceef EF Singapore 2018 · 2018 · B2B SaaSalive

    Building the most candidate-friendly skills assessment platform.

  • Codilityseedcamp Seedcamp 2009 · 2009 · Developer toolsalive

    Codility is an automated tool for assessment of programming skills.

Public money in this direction

US federal grants and open opportunities matched to the concept's terms.

Other concepts in this collection

Fictional concept generated 2026-08-26 by claude-fable-5 from the collection's brief and MarkosWeb data. Treat it as a research prompt, not a plan.