New startup ideas · AI for people who run the AI themselves · Skills, not services
startup concept
Skillproof
Regression testing for the Agent Skills you actually depend on
A test harness for SKILL.md folders: a builder writes assertions against a skill's outputs, then Skillproof replays the suite across Claude, Codex, Copilot, Cursor and Gemini CLI and across model updates, flagging when a skill silently breaks.
- Software subscription
- Consumer
- Rides 'Agent Skills
13
similar startups, last 2 years (42 all-time)
yes
8 matching federal grants and programs
Direction supported by government programs and grants
Test it before you build it
$500 · 3 weeks · 30 prospects
For $500 and 3 weeks, prove that solo skill publishers will prepay $29 to know their SKILL.md still works after the next model update.
Riskiest assumption · A solo skill publisher will pay $29 a month to catch cross-assistant breakage, rather than treating an occasional manual re-run as good enough while waiting for free conformance tooling from the assistant vendors.
1Focus group: who and where
A solo builder who published a SKILL.md skill on GitHub that peers now install into Claude, Cursor or Codex, who pays for Claude Max or Cursor themselves, and who got at least one 'it stopped working' issue after a model release in the last 90 days.
where to find 30 · GitHub code search for repos containing SKILL.md pushed in the last 90 days, sorted by stars (the list); the Claude Developers Discord, in the channels where builders announce skills (the community); r/ClaudeAI and a Show HN post on Hacker News (the channels).
2Sell first, build later
Before any harness exists: a hand-run regression suite on your published skill - 5 golden tasks replayed across Claude, Codex CLI, Cursor and Gemini CLI, pass/fail matrix delivered within 5 days - then a founding slot in the first 20-builder cohort with nightly automated runs starting October 1, 2026.
the ask · $29 one time for the manual run; $29 per month, prepaid first month, for the founding nightly subscription
a real yes · A real yes is $29 charged to their card for the run or the first month; stars, 'this is needed', and offers to beta test for free are noise.
3Small experiments
The first one attacks the riskiest assumption; each ends with a number that says whether to run the next.
1. Sell a manual regression run
$380 · 10 days
Pull 30 publishers from GitHub code search and Discord. Pitch each a $29 prepaid manual run: the founder records 5 golden tasks from their repo, replays them by hand across Claude, Codex CLI, Cursor and Gemini CLI, and delivers a pass/fail matrix in 5 days. Payment before the run; the matrix is the product demo.
keep going if · 6 of 30 pitched publishers pay $29 up front
2. Breakage ritual interviews
$0 · 7 days
Book 12 of the 30 for a 20-minute call, including everyone who declined to pay. Ask for the last silent breakage, how they found out, and what they do before recommending their skill after a model release. Founder runs all calls, notes verbatim.
keep going if · 8 of 12 describe a concrete breakage in the last 90 days and a manual retest ritual they repeat
3. Founding nightly-run presale
$120 · 14 days
One-page site: $29 a month, your suite replayed nightly across the assistants you support, a diff emailed when anything breaks, first cohort of 20. Card checkout for the first month, prepaid. Post it as Show HN, in r/ClaudeAI and in the Discord showcase channel; DM the 30 pitched publishers the link.
keep going if · 10 prepaid first months from roughly 500 visits plus the 30 direct pitches
4Collect a deposit up front
Tesla took $1,000 refundable reservations for the Model 3 and $100 for the Cybertruck before building either: the deposit is the measurement, not the revenue.
$29
per prospect, refundable
how · Prepaid first month through the card checkout on the landing page - this buyer already pays $20 to $200 a month out of pocket for assistants, so a small self-serve charge is the honest instrument, not an LOI. set up: Stripe Checkout ↗
what it reserves · A slot in the 20-builder founding cohort, the $29 price locked for 12 months, and their suite in the nightly run from day one
refund · Fully refunded on request any time before their first nightly diff report is delivered, no questions.
target · 10 prepaid first months from 30 direct pitches plus the public posts, within 21 days
Go: build it if
10 prepaid founding months including 6 or more paid manual runs, and 8 of 12 interviews confirming a recurring manual retest ritual: build the CLI.
Kill: stop if
Fewer than 3 of 30 publishers pay anything, or fewer than 4 of 12 interviews surface a real breakage in the last 90 days: the pain is not worth $29 and a vendor conformance kit will absorb it.
5 Scripts to run itoutreach message, landing copy, deposit terms · click to open
outreach message
You published a skill that other people now run inside Claude and Cursor, which means every model release can silently break it on someone else's machine before you hear about it. I'm testing a regression harness for SKILL.md folders; before writing code I'm hand-running suites for 20 builders: 5 golden tasks from your repo, replayed across Claude, Codex CLI, Cursor and Gemini CLI, pass/fail matrix in 5 days, $29. Worth 20 minutes this week to pick the tasks?
landing page
Know your skill still works before your users do $29 a month, prepaid: your 5 golden tasks replayed nightly across Claude, Codex, Cursor and Gemini CLI, diff emailed on any break Prepay your first month and take one of 20 founding slots
deposit terms
You pay $29 today, your normal first month. It reserves one of 20 founding slots, locks $29 a month for 12 months, and puts your suite in the nightly run starting October 1, 2026. Full refund on request any time before your first diff report arrives.
Would you run this test?
One tap. The yes-share feeds the Demand pillar of this idea's score; nobody sees who answered.
Budgets are out-of-pocket estimates for a team of one to three, US market. Size the deposit to the deal, and check the terms before taking money in a regulated line.
Scorecard
Ranked against every idea in the catalog: trend, demand and 100x potential from the corpus, competition relative to the other ideas. A generated concept has no judges or swipes yet, so its pillars use the data signals only.
61
Idea Score, 0-100 · raw 37.5 x 1.61
Crowded
competition: more crowded than 84% of ideas · headwind x0.58
+4.9
government priorities, secondary (215 matching grants)
Trend
50
Is the wave forming now? 2025-26 entrants vs 2023-24, rounds since 2025, the sector's live-batch direction, the 2026 trend analyst.
- Entrants 2025-26 vs 2023-24 (similar companies)53
- Rounds announced 2025+ in the sector0
- Rounds announced 2025+ matching the idea99
- Sector direction (live batch)50
Demand
65
Does anyone want it? YC's current RFS, companies already paid for something similar, the operator judge, founders' yes-rate in decks, readers who would run the test.
- YC asks for it (current RFS: idea / sector)30
- Someone already pays (similar companies, recent / all-time)100
100x potential
54
Can it return a fund? The venture judge (double weight), market-size and moat axes, neighbours still alive, the technologist judge.
- Neighbours still alive54
Score = 100 x cbrt(Trend x Demand x 100x) x (1 - 0.5 x crowding) + government bonus (max 5), calibrated so the 95th-percentile idea scores 90 (order never changes). A geometric mean: a weak pillar cannot be papered over. Percentiles are among the 382 ideas in the catalog; the terms matched were regression, testing, skills, depend, test, harness, skill, folders.
The concept in full
- What
- A test harness for SKILL.md folders: a builder writes assertions against a skill's outputs, then Skillproof replays the suite across Claude, Codex, Copilot, Cursor and Gemini CLI and across model updates, flagging when a skill silently breaks. Sold self-serve to the individual builders shipping and maintaining their own skills. In the first hour a user points it at a skill folder, records five golden tasks, and gets a pass/fail matrix across the assistants they support.
- Grounded in (2025-2026 signals)
- 'Anthropic's Agent Skills... released as an open standard on December 18, 2025 (agentskills.io), and by June 2026 about 40 products supported it, including Claude, OpenAI Codex, GitHub Copilot, VS Code, Cursor, Gemini CLI and Goose.' Also the Anthropic Economic Index (November 2025 data): 52% augmentation on Claude.ai, the population that iterates on its own tooling.
- What it rides
- Rides 'Agent Skills: launched October 16, 2025, an open standard since December 18': once one skill must work in about 40 host products, hand-testing dies and a per-skill CI layer becomes mandatory.
- Why now
- The standard is eight months old and already has about 40 supporting products as of June 2026; nobody who published a skill in that window has a way to know it still works after the next model release.
- Wedge: first customer and entry point
- First customer: a Claude Max or Cursor subscriber who published a skill their peers now use. Entry point: a CLI plus a $29 a month dashboard that runs their suite nightly and emails a diff, before touching teams or hosted runners.
- Closest real companies, as the generator saw them
- Confident AI (yc W25) does LLM evals for engineering teams building AI products; Skillproof tests the skill artifact itself for the solo builder, per SKILL.md folder, across host assistants. Quantstruct (yc W25) auto-tests stale docs, not skills.
- Main risk
- Assistant vendors bundle a conformance test kit into the open standard's tooling and the standalone harness becomes a feature.
Similar startups in the directory
Companies whose pitch matches most of the concept's terms (regression, testing, skills, depend, test, harness, skill, folders).
Aillis specializes in the development of testing methods for Influenza using artificial intelligence technology.
Jobo is a platform in Ivory Cost offering rapid, reliable, and automated temporary employment services. They use AI to preselect candidates within 24 hours and ensure all temporary workers are trained before starting their roles.
Simulation & Evaluation that scales voice and chat AI agents
Operator of an online learning platform intended to improve students' academic skills. The company provides an online diagnostic test to analyze student performance and provide information about their academic abilities as well as collaboration tools to allow students to work together on projects and assignments, enabling students to learn at their own pace with the opportunity to interact with other students and instructors to get feedback on their work.
The Computer Science Proficiency Assessment (CSPA™) is a…
Building the most candidate-friendly skills assessment platform.
Codility is an automated tool for assessment of programming skills.
Game-based coding for kids ages 8–14.
Robots for high skilled labor powering AI infrastructure
Catch and fix AI agent regressions before they reach production
Autonomous Hardware Testing
Ship more code with confidence
Public money in this direction
US federal grants and open opportunities matched to the concept's terms.
NIH / NIBIB · SBIR phase I · $264K
National Science Foundation · SBIR Phase II · $1M
NIH / NIMHD · SBIR phase II · $1M
NIH / NIMHD · SBIR phase II · $1M
NIH / NIMHD · SBIR phase I · $350K
NIH / NIA · STTR phase I · $498K
NIH / NIA · SBIR phase I · $349K
NIH / NIBIB · STTR phase I · $320K
Other concepts in this collection
- ProvenaryScan third-party skills and MCP servers before you let them touch your data
- LedgerkitVersioned skill packs that make a solo CPA's assistant work like a tax practice
- VendfoldLicensing, signing and auto-update infrastructure for people who sell Agent Skills
- PackroomOne shared skill library for a team where everyone runs their own agent
- TokentabPer-skill cost, routing and drift telemetry for the person who runs AI all day
- ThreadkeepA memory vault you own that every assistant you run can read
- RelayfileHand a running task from Claude Code to Codex without losing state
- MeterhouseOne budget, meter and kill switch for every agent you run
- AttestlyAudit trail and approval inbox for the agents you run at work
- SkillvaneVersion control and regression tests for the skills your agents load
- CrewlineA shared board where each teammate's agents pick up each other's work
- WardkeySecurity scanner that finds and fixes exposed keys in vibe-coded apps
- StillupUptime and error monitoring that answers in fix prompts, not stack traces
- CopystoneAutomatic backups and one-click restore for apps built without engineers
- GroundskeepMonthly maintenance for shipped vibe-coded apps, applied as reviewable patches
- TillhousePayments, sales tax and refunds as one drop-in for non-developer founders
- SpendgateMeter, cap and route the AI spend inside apps vibe coders shipped
- DryloopRehearsal mode for the automations a small business owner builds alone
- MeterlyOne metered key with spend caps for every AI step you run
- FlowmedicWatches your automations, explains failures in plain English, proposes the fix
- ScrubdeckA data-cleaning step any workflow can call, with rules the owner keeps
- OpshandTurns your written SOPs into versioned Agent Skills with tests included
- CrewtraceShared visibility when five people at one business each run their own automations
- VeraciteCitation verification and AI work records for solo attorneys who draft with Claude
- TickstoneTurns a solo CPA's AI sessions into reviewable workpapers with tickmarks and source trails
- ChartproofA verification layer for physicians who use AI on clinical notes under their own license
- CoverlensPolicy-form verification for independent insurance agents who quote with AI
- MethodkitSolo consultants package their methodology as versioned Agent Skills they own and resell
- AttestrailTamper-evident logs of every AI action, built for licensed professionals' liability files
- ScrublineLocal redaction proxy that makes your personal AI accounts safe for work data
- StipendlyTurn personal Claude Max and ChatGPT Pro seats into managed employer stipends
- TollgateA policy gateway between your assistant and every MCP server it touches
- SkillvetScan, pin and approve Agent Skills before they touch company data
- DaylightSelf-serve shadow AI registry and policy for companies with no security team
- LedgerlineRightsizing dashboard for everyone paying for AI out of their own pocket
- SwitchyardOne metered endpoint with routing, fallback and per-person caps for tiny teams
- HearthmeterUsage budgets and one bill for the household that shares AI plans
- SeatcaseMeasures who on your team earns a Max seat and who wastes one
- TokencairnProfiler that shows what each installed skill and MCP server really costs
- FusegateBudget caps, fallback and kill switches for automations you run yourself
- SkillbenchRegression testing for Agent Skills before every model and skill update
- CitelockVerifies every citation in AI-drafted work before a licensed professional signs it
- MiddlegateA local gateway where you set the rules for what your MCP servers can do
- DriftwatchCatches output drift in the automations small operators wired themselves
- ShipcheckPre-launch review gates non-technical builders run on their own vibe-coded apps
- TracelineA claim-level provenance trail for every number in an AI-assisted report
- DrillyardScored practice repos where you learn to drive coding agents well
- PassrateA proctored AI operation exam scored from your real agent transcripts
- PatchcraftDebugging drills that teach non-technical builders to maintain what they vibe coded
- SkillsmithA workshop for writing, testing and versioning Agent Skills that actually hold up
- TickmarkSynthetic client caseloads where CPAs drill AI-assisted work before trying it on real clients
- PostgameAn MCP server that scores your own agent sessions and drills your weakest habits
- CitegridEvery number in your published research links to a source snapshot you verified
- MnemosYour research corpus as a private MCP server every assistant can query
- MeterlineModel routing and cost accounting for one person's AI research pipeline
- SkillcaskVersion, test, and sell your expertise as licensed Agent Skills
- StackfeedA personal data pipeline that repairs itself when sources change
- ClaimboardA shared evidence ledger for small teams where everyone runs their own agent
Fictional concept generated 2026-08-26 by claude-fable-5 from the collection's brief and MarkosWeb data. Treat it as a research prompt, not a plan.