New startup ideas · AI for people who run the AI themselves · Evaluation and trust the user controls
startup concept
Skillbench
Regression testing for Agent Skills before every model and skill update
A hosted test harness for Agent Skills.
- Software subscription
- Small business
- Agent Skills
54
similar startups, last 2 years (143 all-time)
yes
8 matching federal grants and programs
Direction supported by government programs and grants
Test it before you build it
$270 · 3 weeks · 25 prospects
For $270 and 3 weeks, prove that Agent Skills authors will pay cash today for regression tests of their SKILL.md, starting with a $49 manual run.
Riskiest assumption · Individual skill authors who distribute skills to others exist in paying numbers and will pay $30 a month for third-party regression testing, rather than eyeballing a few prompts after each update and waiting for agentskills.io or a major runtime to ship a free conformance suite.
1Focus group: who and where
An individual author who publishes or sells Agent Skills - a public repo with a SKILL.md and real users - pays $100 to $200 a month for a Max-tier seat, and watched a skill misbehave after a model or runtime update in the past 90 days.
where to find 25 · GitHub code search for repos containing SKILL.md sorted by stars and recent activity (a live directory of exactly these authors), the official Claude Developers Discord, skill-sharing threads on r/ClaudeAI, and author pages on agentskills.io.
2Sell first, build later
Today: a $49 manual regression run - ten fixture tasks built from your transcripts, replayed against Sonnet 5 and Haiku 4.5, line-by-line diff report in 48 hours. By October 15, 2026: the hosted harness that reruns your pinned suite automatically on every model version and skill edit.
the ask · $49 per manual run now; $99 founding year for the hosted harness (then $30 a month)
a real yes · A real yes is $49 or $99 actually charged. 'I'd try the free tier', GitHub stars, and offers to beta test for free do not count.
3Small experiments
The first one attacks the riskiest assumption; each ends with a number that says whether to run the next.
1. Fifteen author regression interviews
$0 · 10 days
Pull the 40 most-starred, recently updated SKILL.md repos from GitHub code search and contact the authors through repo issues, Discord, or their listed contact. Book 15 twenty-minute calls, get the story of their last model-update breakage in specifics, and close each call with a money ask: $49 for a manual regression run of their skill, ten fixtures, two models, results in 48 hours.
keep going if · 9 of 15 shipped a regression after a model update in the past 90 days, and 5 of 15 pay the $49 on the call
2. Paid manual regression runs
$150 · 12 days
Deliver every purchased run by hand: write ten fixture tasks from the author's own transcripts and examples, replay the skill against Sonnet 5 and Haiku 4.5 via the API (about $150 of model spend across runs), and send a line-by-line diff report. Manual delivery is the right first artifact here because the product's entire value is this report, and producing it with the models directly is faster than any prototype harness.
keep going if · 5 of 20 authors asked buy the $49 run, and 3 of 5 buyers say they would rerun it before every model update
3. Founding-author pre-order page
$120 · 14 days
Put up a one-page site with a redacted diff report from a real run, the $99 founding year (regular $360 at $30 a month), delivery October 15, 2026, and card checkout. Send it to every interviewee and buyer, and post it once in the Claude Developers Discord and once on r/ClaudeAI.
keep going if · 8 of 25 authors asked pay the $99 founding year
4Collect a deposit up front
Tesla took $1,000 refundable reservations for the Model 3 and $100 for the Cybertruck before building either: the deposit is the measurement, not the revenue.
$99
per prospect, refundable
how · Two rungs of real money by card checkout: the $49 manual run (money moved today for the exact artifact the product automates) and the $99 refundable founding-year pre-order. Card checkout fits a solo author buying a sub-$100 developer tool online with nobody else to sign off. set up: Stripe Checkout ↗
what it reserves · The founding price of $99 a year for life of the subscription (versus $360), a slot in the first cohort on October 15, 2026, and their fixture suite migrated into the harness for free
refund · The $99 is fully refundable by email until launch on October 15, 2026 and for 30 days after; delivered $49 runs are not refundable.
target · 5 paid $49 runs and 8 paid $99 pre-orders from 25 authors within 21 days
Go: build it if
5 or more $49 runs sold, 8 or more $99 pre-orders, and 3 of 5 run-buyers saying they would rerun before every model update: build the harness.
Kill: stop if
Fewer than 3 paid runs after 25 asks, or fewer than 5 of 15 interviewed authors report any regression in the past 90 days: stop, the pain is too rare to beat a free native conformance suite when one ships.
5 Scripts to run itoutreach message, landing copy, deposit terms · click to open
outreach message
Your skill has real users on GitHub, which means every model update can silently break it for all of them before you notice. I'm building Skillbench: point it at your SKILL.md, record ten fixture tasks, and it replays the suite against every new model version, flagging drift line by line. Before writing the product I'm doing paid manual runs: $49, ten fixtures from your own transcripts, two models, diff report back in 48 hours. Got 20 minutes this week to scope yours?
landing page
Know your skill still works before your users find out $49 per manual regression run today; $99 founding year for the hosted harness (then $360) Buy a run or reserve your founding year - $99, refundable until launch on October 15, 2026
deposit terms
Your $99 buys the founding year of Skillbench, live October 15, 2026, and locks your price at $99 a year instead of $360; we also migrate your fixture suite in for free. Fully refundable by email until launch and for 30 days after. Manual $49 runs are delivered within 48 hours and are not refundable once delivered.
Would you run this test?
One tap. The yes-share feeds the Demand pillar of this idea's score; nobody sees who answered.
Budgets are out-of-pocket estimates for a team of one to three, US market. Size the deposit to the deal, and check the terms before taking money in a regulated line.
Scorecard
Ranked against every idea in the catalog: trend, demand and 100x potential from the corpus, competition relative to the other ideas. A generated concept has no judges or swipes yet, so its pillars use the data signals only.
59
Idea Score, 0-100 · raw 36.5 x 1.61
Crowded
competition: more crowded than 98% of ideas · headwind x0.51
+5.0
government priorities, secondary (667 matching grants)
Trend
55
Is the wave forming now? 2025-26 entrants vs 2023-24, rounds since 2025, the sector's live-batch direction, the 2026 trend analyst.
- Entrants 2025-26 vs 2023-24 (similar companies)70
- Rounds announced 2025+ in the sector0
- Rounds announced 2025+ matching the idea99
- Sector direction (live batch)50
Demand
65
Does anyone want it? YC's current RFS, companies already paid for something similar, the operator judge, founders' yes-rate in decks, readers who would run the test.
- YC asks for it (current RFS: idea / sector)30
- Someone already pays (similar companies, recent / all-time)100
100x potential
67
Can it return a fund? The venture judge (double weight), market-size and moat axes, neighbours still alive, the technologist judge.
- Neighbours still alive67
Score = 100 x cbrt(Trend x Demand x 100x) x (1 - 0.5 x crowding) + government bonus (max 5), calibrated so the 95th-percentile idea scores 90 (order never changes). A geometric mean: a weak pillar cannot be papered over. Percentiles are among the 382 ideas in the catalog; the terms matched were regression, testing, skills, model, skill, hosted, test, harness.
The concept in full
- What
- A hosted test harness for Agent Skills. A builder points it at a SKILL.md folder, records fixture tasks with expected outputs, and Skillbench replays the suite against new model versions and skill edits, flagging behavior drift line by line. In the first hour a user imports one skill, auto-generates ten fixture cases from past transcripts, and gets a pass or fail report they can pin to a version.
- Grounded in (2025-2026 signals)
- 'Agent Skills package procedural knowledge as folders with a SKILL.md' launched October 16, 2025, 'released as an open standard on December 18, 2025 (agentskills.io), and by June 2026 about 40 products supported it, including Claude, OpenAI Codex, GitHub Copilot, VS Code, Cursor, Gemini CLI and Goose'. Claude Max launched April 9, 2025 at $100 and $200 a month for people who run the model all day.
- What it rides
- Agent Skills: launched October 16, 2025, an open standard since December 18. Skills are now versioned artifacts running across many assistants, so a change in the skill or the underlying model silently changes results everywhere at once, and nothing in the standard tests for that.
- Why now
- About 40 products supported the Agent Skills standard by June 2026, eight months after the October 16, 2025 launch, so one skill now runs on many runtimes and models with no shared test surface; the people maintaining those skills already pay $100 to $200 a month for Max-tier seats and will pay $30 more to stop shipping regressions.
- Wedge: first customer and entry point
- Individual skill authors selling or sharing skills built with the skill-creator skill; entry point is a free tier that tests one skill against two models, converting authors whose skill broke on a model update last month.
- Closest real companies, as the generator saw them
- Confident AI (yc W25) is an LLM eval and observability platform aimed at engineering teams instrumenting their own apps; Skillbench tests the SKILL.md artifact itself across third-party runtimes for a single author. Quantstruct (yc W25) tests stale product docs, not executable skills.
- Main risk
- The agentskills.io standard body or a major runtime ships a built-in conformance test suite and authors never look for a third-party one.
Similar startups in the directory
Companies whose pitch matches most of the concept's terms (regression, testing, skills, model, skill, hosted, test, harness).
Aillis specializes in the development of testing methods for Influenza using artificial intelligence technology.
Catch and fix AI agent regressions before they reach production
Robots for high skilled labor powering AI infrastructure
Monitoring and learning layer for long-running agents
Harness Engineering Infrastructure that developers love.
Jobo is a platform in Ivory Cost offering rapid, reliable, and automated temporary employment services. They use AI to preselect candidates within 24 hours and ensure all temporary workers are trained before starting their roles.
Personalized prevention using genetics and AI
Operator of an online learning platform intended to improve students' academic skills. The company provides an online diagnostic test to analyze student performance and provide information about their academic abilities as well as collaboration tools to allow students to work together on projects and assignments, enabling students to learn at their own pace with the opportunity to interact with other students and instructors to get feedback on their work.
The AI Cowork for big data
The Computer Science Proficiency Assessment (CSPA™) is a…
Building the most candidate-friendly skills assessment platform.
Codility is an automated tool for assessment of programming skills.
Public money in this direction
US federal grants and open opportunities matched to the concept's terms.
NIH / NIA · SBIR phase I · $349K
NIH / NIA · STTR phase I · $498K
NIH / NIBIB · STTR phase I · $320K
National Science Foundation · SBIR Fast-Track · $2M
NIH / NIAMS · STTR phase I · $314K
National Science Foundation · I-Corps · $50K
NIH / NIDDK · SBIR phase II · $905K
NIH / NHLBI · SBIR phase I · $599K
Other concepts in this collection
- SkillproofRegression testing for the Agent Skills you actually depend on
- ProvenaryScan third-party skills and MCP servers before you let them touch your data
- LedgerkitVersioned skill packs that make a solo CPA's assistant work like a tax practice
- VendfoldLicensing, signing and auto-update infrastructure for people who sell Agent Skills
- PackroomOne shared skill library for a team where everyone runs their own agent
- TokentabPer-skill cost, routing and drift telemetry for the person who runs AI all day
- ThreadkeepA memory vault you own that every assistant you run can read
- RelayfileHand a running task from Claude Code to Codex without losing state
- MeterhouseOne budget, meter and kill switch for every agent you run
- AttestlyAudit trail and approval inbox for the agents you run at work
- SkillvaneVersion control and regression tests for the skills your agents load
- CrewlineA shared board where each teammate's agents pick up each other's work
- WardkeySecurity scanner that finds and fixes exposed keys in vibe-coded apps
- StillupUptime and error monitoring that answers in fix prompts, not stack traces
- CopystoneAutomatic backups and one-click restore for apps built without engineers
- GroundskeepMonthly maintenance for shipped vibe-coded apps, applied as reviewable patches
- TillhousePayments, sales tax and refunds as one drop-in for non-developer founders
- SpendgateMeter, cap and route the AI spend inside apps vibe coders shipped
- DryloopRehearsal mode for the automations a small business owner builds alone
- MeterlyOne metered key with spend caps for every AI step you run
- FlowmedicWatches your automations, explains failures in plain English, proposes the fix
- ScrubdeckA data-cleaning step any workflow can call, with rules the owner keeps
- OpshandTurns your written SOPs into versioned Agent Skills with tests included
- CrewtraceShared visibility when five people at one business each run their own automations
- VeraciteCitation verification and AI work records for solo attorneys who draft with Claude
- TickstoneTurns a solo CPA's AI sessions into reviewable workpapers with tickmarks and source trails
- ChartproofA verification layer for physicians who use AI on clinical notes under their own license
- CoverlensPolicy-form verification for independent insurance agents who quote with AI
- MethodkitSolo consultants package their methodology as versioned Agent Skills they own and resell
- AttestrailTamper-evident logs of every AI action, built for licensed professionals' liability files
- ScrublineLocal redaction proxy that makes your personal AI accounts safe for work data
- StipendlyTurn personal Claude Max and ChatGPT Pro seats into managed employer stipends
- TollgateA policy gateway between your assistant and every MCP server it touches
- SkillvetScan, pin and approve Agent Skills before they touch company data
- DaylightSelf-serve shadow AI registry and policy for companies with no security team
- LedgerlineRightsizing dashboard for everyone paying for AI out of their own pocket
- SwitchyardOne metered endpoint with routing, fallback and per-person caps for tiny teams
- HearthmeterUsage budgets and one bill for the household that shares AI plans
- SeatcaseMeasures who on your team earns a Max seat and who wastes one
- TokencairnProfiler that shows what each installed skill and MCP server really costs
- FusegateBudget caps, fallback and kill switches for automations you run yourself
- CitelockVerifies every citation in AI-drafted work before a licensed professional signs it
- MiddlegateA local gateway where you set the rules for what your MCP servers can do
- DriftwatchCatches output drift in the automations small operators wired themselves
- ShipcheckPre-launch review gates non-technical builders run on their own vibe-coded apps
- TracelineA claim-level provenance trail for every number in an AI-assisted report
- DrillyardScored practice repos where you learn to drive coding agents well
- PassrateA proctored AI operation exam scored from your real agent transcripts
- PatchcraftDebugging drills that teach non-technical builders to maintain what they vibe coded
- SkillsmithA workshop for writing, testing and versioning Agent Skills that actually hold up
- TickmarkSynthetic client caseloads where CPAs drill AI-assisted work before trying it on real clients
- PostgameAn MCP server that scores your own agent sessions and drills your weakest habits
- CitegridEvery number in your published research links to a source snapshot you verified
- MnemosYour research corpus as a private MCP server every assistant can query
- MeterlineModel routing and cost accounting for one person's AI research pipeline
- SkillcaskVersion, test, and sell your expertise as licensed Agent Skills
- StackfeedA personal data pipeline that repairs itself when sources change
- ClaimboardA shared evidence ledger for small teams where everyone runs their own agent
Fictional concept generated 2026-08-26 by claude-fable-5 from the collection's brief and MarkosWeb data. Treat it as a research prompt, not a plan.