New startup ideas · AI for people who run the AI themselves · Skills, not services

startup concept

Skillproof

Regression testing for the Agent Skills you actually depend on

A test harness for SKILL.md folders: a builder writes assertions against a skill's outputs, then Skillproof replays the suite across Claude, Codex, Copilot, Cursor and Gemini CLI and across model updates, flagging when a skill silently breaks.

13

similar startups, last 2 years (42 all-time)

yes

8 matching federal grants and programs

Direction supported by government programs and grants

Test it before you build it

$500 · 3 weeks · 30 prospects

For $500 and 3 weeks, prove that solo skill publishers will prepay $29 to know their SKILL.md still works after the next model update.

Riskiest assumption · A solo skill publisher will pay $29 a month to catch cross-assistant breakage, rather than treating an occasional manual re-run as good enough while waiting for free conformance tooling from the assistant vendors.

1Focus group: who and where

A solo builder who published a SKILL.md skill on GitHub that peers now install into Claude, Cursor or Codex, who pays for Claude Max or Cursor themselves, and who got at least one 'it stopped working' issue after a model release in the last 90 days.

where to find 30 · GitHub code search for repos containing SKILL.md pushed in the last 90 days, sorted by stars (the list); the Claude Developers Discord, in the channels where builders announce skills (the community); r/ClaudeAI and a Show HN post on Hacker News (the channels).

2Sell first, build later

Before any harness exists: a hand-run regression suite on your published skill - 5 golden tasks replayed across Claude, Codex CLI, Cursor and Gemini CLI, pass/fail matrix delivered within 5 days - then a founding slot in the first 20-builder cohort with nightly automated runs starting October 1, 2026.

the ask · $29 one time for the manual run; $29 per month, prepaid first month, for the founding nightly subscription

a real yes · A real yes is $29 charged to their card for the run or the first month; stars, 'this is needed', and offers to beta test for free are noise.

3Small experiments

The first one attacks the riskiest assumption; each ends with a number that says whether to run the next.

  1. 1. Sell a manual regression run

    $380 · 10 days

    Pull 30 publishers from GitHub code search and Discord. Pitch each a $29 prepaid manual run: the founder records 5 golden tasks from their repo, replays them by hand across Claude, Codex CLI, Cursor and Gemini CLI, and delivers a pass/fail matrix in 5 days. Payment before the run; the matrix is the product demo.

    keep going if · 6 of 30 pitched publishers pay $29 up front

  2. 2. Breakage ritual interviews

    $0 · 7 days

    Book 12 of the 30 for a 20-minute call, including everyone who declined to pay. Ask for the last silent breakage, how they found out, and what they do before recommending their skill after a model release. Founder runs all calls, notes verbatim.

    keep going if · 8 of 12 describe a concrete breakage in the last 90 days and a manual retest ritual they repeat

  3. 3. Founding nightly-run presale

    $120 · 14 days

    One-page site: $29 a month, your suite replayed nightly across the assistants you support, a diff emailed when anything breaks, first cohort of 20. Card checkout for the first month, prepaid. Post it as Show HN, in r/ClaudeAI and in the Discord showcase channel; DM the 30 pitched publishers the link.

    keep going if · 10 prepaid first months from roughly 500 visits plus the 30 direct pitches

4Collect a deposit up front

Tesla took $1,000 refundable reservations for the Model 3 and $100 for the Cybertruck before building either: the deposit is the measurement, not the revenue.

$29

per prospect, refundable

how · Prepaid first month through the card checkout on the landing page - this buyer already pays $20 to $200 a month out of pocket for assistants, so a small self-serve charge is the honest instrument, not an LOI. set up: Stripe Checkout

what it reserves · A slot in the 20-builder founding cohort, the $29 price locked for 12 months, and their suite in the nightly run from day one

refund · Fully refunded on request any time before their first nightly diff report is delivered, no questions.

target · 10 prepaid first months from 30 direct pitches plus the public posts, within 21 days

Go: build it if

10 prepaid founding months including 6 or more paid manual runs, and 8 of 12 interviews confirming a recurring manual retest ritual: build the CLI.

Kill: stop if

Fewer than 3 of 30 publishers pay anything, or fewer than 4 of 12 interviews surface a real breakage in the last 90 days: the pain is not worth $29 and a vendor conformance kit will absorb it.

5 Scripts to run itoutreach message, landing copy, deposit terms · click to open

outreach message

You published a skill that other people now run inside Claude and Cursor, which means every model release can silently break it on someone else's machine before you hear about it. I'm testing a regression harness for SKILL.md folders; before writing code I'm hand-running suites for 20 builders: 5 golden tasks from your repo, replayed across Claude, Codex CLI, Cursor and Gemini CLI, pass/fail matrix in 5 days, $29. Worth 20 minutes this week to pick the tasks?

landing page

Know your skill still works before your users do $29 a month, prepaid: your 5 golden tasks replayed nightly across Claude, Codex, Cursor and Gemini CLI, diff emailed on any break Prepay your first month and take one of 20 founding slots

deposit terms

You pay $29 today, your normal first month. It reserves one of 20 founding slots, locks $29 a month for 12 months, and puts your suite in the nightly run starting October 1, 2026. Full refund on request any time before your first diff report arrives.

Would you run this test?

One tap. The yes-share feeds the Demand pillar of this idea's score; nobody sees who answered.

Budgets are out-of-pocket estimates for a team of one to three, US market. Size the deposit to the deal, and check the terms before taking money in a regulated line.

Scorecard

Ranked against every idea in the catalog: trend, demand and 100x potential from the corpus, competition relative to the other ideas. A generated concept has no judges or swipes yet, so its pillars use the data signals only.

61

Idea Score, 0-100 · raw 37.5 x 1.61

Crowded

competition: more crowded than 84% of ideas · headwind x0.58

+4.9

government priorities, secondary (215 matching grants)

Trend

50

Is the wave forming now? 2025-26 entrants vs 2023-24, rounds since 2025, the sector's live-batch direction, the 2026 trend analyst.

  • Entrants 2025-26 vs 2023-24 (similar companies)53
  • Rounds announced 2025+ in the sector0
  • Rounds announced 2025+ matching the idea99
  • Sector direction (live batch)50

Demand

65

Does anyone want it? YC's current RFS, companies already paid for something similar, the operator judge, founders' yes-rate in decks, readers who would run the test.

  • YC asks for it (current RFS: idea / sector)30
  • Someone already pays (similar companies, recent / all-time)100

100x potential

54

Can it return a fund? The venture judge (double weight), market-size and moat axes, neighbours still alive, the technologist judge.

  • Neighbours still alive54

Score = 100 x cbrt(Trend x Demand x 100x) x (1 - 0.5 x crowding) + government bonus (max 5), calibrated so the 95th-percentile idea scores 90 (order never changes). A geometric mean: a weak pillar cannot be papered over. Percentiles are among the 382 ideas in the catalog; the terms matched were regression, testing, skills, depend, test, harness, skill, folders.

The concept in full

What
A test harness for SKILL.md folders: a builder writes assertions against a skill's outputs, then Skillproof replays the suite across Claude, Codex, Copilot, Cursor and Gemini CLI and across model updates, flagging when a skill silently breaks. Sold self-serve to the individual builders shipping and maintaining their own skills. In the first hour a user points it at a skill folder, records five golden tasks, and gets a pass/fail matrix across the assistants they support.
Grounded in (2025-2026 signals)
'Anthropic's Agent Skills... released as an open standard on December 18, 2025 (agentskills.io), and by June 2026 about 40 products supported it, including Claude, OpenAI Codex, GitHub Copilot, VS Code, Cursor, Gemini CLI and Goose.' Also the Anthropic Economic Index (November 2025 data): 52% augmentation on Claude.ai, the population that iterates on its own tooling.
What it rides
Rides 'Agent Skills: launched October 16, 2025, an open standard since December 18': once one skill must work in about 40 host products, hand-testing dies and a per-skill CI layer becomes mandatory.
Why now
The standard is eight months old and already has about 40 supporting products as of June 2026; nobody who published a skill in that window has a way to know it still works after the next model release.
Wedge: first customer and entry point
First customer: a Claude Max or Cursor subscriber who published a skill their peers now use. Entry point: a CLI plus a $29 a month dashboard that runs their suite nightly and emails a diff, before touching teams or hosted runners.
Closest real companies, as the generator saw them
Confident AI (yc W25) does LLM evals for engineering teams building AI products; Skillproof tests the skill artifact itself for the solo builder, per SKILL.md folder, across host assistants. Quantstruct (yc W25) auto-tests stale docs, not skills.
Main risk
Assistant vendors bundle a conformance test kit into the open standard's tooling and the standalone harness becomes a feature.

Similar startups in the directory

Companies whose pitch matches most of the concept's terms (regression, testing, skills, depend, test, harness, skill, folders).

  • アイリス株式会社plugandplay · Healthcare and bioalive

    Aillis specializes in the development of testing methods for Influenza using artificial intelligence technology.

  • Jobo Interimplugandplay PnP 2024 · 2024 · Commerce and marketplacesalive

    Jobo is a platform in Ivory Cost offering rapid, reliable, and automated temporary employment services. They use AI to preselect candidates within 24 hours and ensure all temporary workers are trained before starting their roles.

  • Covalyc S24 · 2024 · Agent infrastructurealive

    Simulation & Evaluation that scales voice and chat AI agents

  • Kings500global 500G Georgia 3 · 2023 · Consumeralive

    Operator of an online learning platform intended to improve students' academic skills. The company provides an online diagnostic test to analyze student performance and provide information about their academic abilities as well as collaboration tools to allow students to work together on projects and assignments, enabling students to learn at their own pace with the opportunity to interact with other students and instructors to get feedback on their work.

  • CSPAyc S18 · 2018 · Developer toolsacquired

    The Computer Science Proficiency Assessment (CSPA™) is a…

  • Adafaceef EF Singapore 2018 · 2018 · B2B SaaSalive

    Building the most candidate-friendly skills assessment platform.

  • Codilityseedcamp Seedcamp 2009 · 2009 · Developer toolsalive

    Codility is an automated tool for assessment of programming skills.

  • Code Kingdomsef · Consumeralive

    Game-based coding for kids ages 8–14.

  • Shepherd Roboticsyc F26 · 2026 · Robotics and physical worldalive

    Robots for high skilled labor powering AI infrastructure

  • Mirrorsyc F26 · 2026 · Agent infrastructurealive

    Catch and fix AI agent regressions before they reach production

  • Hilstartyc S26 · 2026 · Agent infrastructurealive

    Autonomous Hardware Testing

  • Alchemizeyc X26 · 2026 · Developer toolsalive

    Ship more code with confidence

Public money in this direction

US federal grants and open opportunities matched to the concept's terms.

Other concepts in this collection

Fictional concept generated 2026-08-26 by claude-fable-5 from the collection's brief and MarkosWeb data. Treat it as a research prompt, not a plan.