New startup ideas · AI for people who run the AI themselves · Workbenches for people who run many agents
startup concept
Skillvane
Version control and regression tests for the skills your agents load
Skillvane is a registry and CI service for SKILL.md folders and MCP server configs: it versions them, runs them against Claude, Codex, Copilot and Cursor on every model release, flags behavioral drift, and pins or rolls back with one call.
- Infrastructure and APIs
- Small business
- 'Agent Skills
11
similar startups, last 2 years (31 all-time)
yes
8 matching federal grants and programs
Direction supported by government programs and grants
Test it before you build it
$400 · 3 weeks · 25 prospects
For $400 and three weeks, prove that teams sharing a Claude Code skills repo will pay $149 up front for a drift audit and prepay a founding month of automated regression runs.
Riskiest assumption · Small teams sharing a Claude Code skills repo lose enough hours to silent skill breakage after model releases that they will pay up front for regression runs instead of eyeballing failures or waiting for the assistant vendors to build testing in
1Focus group: who and where
The engineering lead of a 2-10 person team that keeps a shared .claude/skills repo their whole team loads into Claude Code, Cursor or Copilot, who spent hours this month figuring out why a skill behaved differently after a model release
where to find 25 · GitHub code search for public repos containing SKILL.md and .claude/skills folders, contacting the maintainers directly; the claude-code channels of the official Anthropic Discord and r/ClaudeAI threads about skills breaking after releases; AI Tinkerers monthly meetups in SF, NYC or Seattle for in-person pitches
2Sell first, build later
A $149 drift audit of the team's skills repo across four assistants with a written diff report in 5 business days, then a founding plan that runs the same tests automatically on every model release for up to 15 skills, starting October 2026
the ask · $149 one-time for the audit; $50 per month prepaid for the founding plan, locked for a year
a real yes · A real yes is $149 charged before any work starts, or a $50 founding month prepaid on the readout call; GitHub stars, 'this should exist' replies and requests for a free beta are not
3Small experiments
The first one attacks the riskiest assumption; each ends with a number that says whether to run the next.
1. Maintainer drift interviews
$0 · 7 days
Pull 25 active public repos with SKILL.md folders via GitHub code search, plus posters in the Anthropic Discord complaining about skill behavior changing. DM or email the maintainers referencing their specific repo and book 12 twenty-minute calls. Ask when a model release last broke a skill, how they found out, how long it cost them, and what they do today before pinning a model version.
keep going if · 7 of 12 describe a concrete skill breakage after a model release in the last 90 days that cost them an hour or more
2. Priced audit page
$60 · 10 days
One-page site: a $149 drift audit of your skills repo, run across Claude, Codex, Copilot and Cursor on the current and previous model versions, exact diffs in a written report within 5 business days. Pitch it directly to the 12 interviewees and 8 more maintainers from the GitHub list; checkout on the page. Money charged is the only metric.
keep going if · 5 of 20 directly pitched teams pay $149
3. Deliver audits, upsell founding plan
$340 · 10 days
For each paying team, run their skills by hand through the assistant CLIs against both model versions, write the diff report, and deliver it on a 30-minute readout call. At the end of the call offer the founding plan: automated runs on every model release, up to 15 skills, $50 a month prepaid, price locked for a year, with the $149 credited against it.
keep going if · 3 of 5 audit buyers prepay the $50 founding month on the readout call
4Collect a deposit up front
Tesla took $1,000 refundable reservations for the Model 3 and $100 for the Cybertruck before building either: the deposit is the measurement, not the revenue.
$149
per prospect, refundable
how · The audit fee itself, invoiced up front through the checkout page before any work begins, because a small dev team will pay for a concrete deliverable long before they subscribe to infrastructure; the $149 credits toward the first founding month set up: Stripe Invoicing ↗
what it reserves · An audit slot in the next 5 business days, the written drift report, a 30-minute readout call, and the $50 founding price locked for a year if they subscribe
refund · Refunded in full if the report is more than 2 business days late or the team cancels before repo access is granted
target · 5 paid audits from 20 direct pitches within 21 days
Go: build it if
7 of 12 maintainers report a recent costly breakage, 5 of 20 pay $149, and 3 of 5 buyers prepay the $50 founding month: build the CI service
Kill: stop if
Fewer than 4 of 12 can name a real breakage, or 1 or fewer audits sell after 20 direct pitches: teams eyeball drift for free and wait for Anthropic to ship this, stop
5 Scripts to run itoutreach message, landing copy, deposit terms · click to open
outreach message
Your team keeps shared skills in a repo, and every Claude, Codex or Cursor model release can silently change how those skills behave. I run your skills against the current and previous model versions across all four assistants and send you a drift report with exact diffs: $149, delivered within five business days. Before that, I'd value 20 minutes on how the last release treated your repo, whatever you decide about the audit. Do you have a slot this week?
landing page
Did the last model release quietly break your agent skills? $149 buys a full drift audit of your skills repo across four assistants, report in 5 business days. Book your audit - refunded in full if the report is late.
deposit terms
$149 is charged today and buys a drift audit of up to 15 skills, with a written diff report and a 30-minute readout within 5 business days of granting repo access. It credits in full toward your first founding-plan month if you subscribe. Refunded entirely if the report is more than 2 business days late or you cancel before we get repo access.
Would you run this test?
One tap. The yes-share feeds the Demand pillar of this idea's score; nobody sees who answered.
Budgets are out-of-pocket estimates for a team of one to three, US market. Size the deposit to the deal, and check the terms before taking money in a regulated line.
Scorecard
Ranked against every idea in the catalog: trend, demand and 100x potential from the corpus, competition relative to the other ideas. A generated concept has no judges or swipes yet, so its pillars use the data signals only.
62
Idea Score, 0-100 · raw 38.2 x 1.61
Crowded
competition: more crowded than 80% of ideas · headwind x0.60
+4.8
government priorities, secondary (154 matching grants)
Trend
51
Is the wave forming now? 2025-26 entrants vs 2023-24, rounds since 2025, the sector's live-batch direction, the 2026 trend analyst.
- Entrants 2025-26 vs 2023-24 (similar companies)54
- Rounds announced 2025+ in the sector0
- Rounds announced 2025+ matching the idea99
- Sector direction (live batch)50
Demand
65
Does anyone want it? YC's current RFS, companies already paid for something similar, the operator judge, founders' yes-rate in decks, readers who would run the test.
- YC asks for it (current RFS: idea / sector)30
- Someone already pays (similar companies, recent / all-time)100
100x potential
53
Can it return a fund? The venture judge (double weight), market-size and moat axes, neighbours still alive, the technologist judge.
- Neighbours still alive53
Score = 100 x cbrt(Trend x Demand x 100x) x (1 - 0.5 x crowding) + government bonus (max 5), calibrated so the 95th-percentile idea scores 90 (order never changes). A geometric mean: a weak pillar cannot be papered over. Percentiles are among the 382 ideas in the catalog; the terms matched were version, control, regression, tests, skills, load, registry, skill.
The concept in full
- What
- Skillvane is a registry and CI service for SKILL.md folders and MCP server configs: it versions them, runs them against Claude, Codex, Copilot and Cursor on every model release, flags behavioral drift, and pins or rolls back with one call. In the first hour a user points it at their skills folder, gets everything versioned, and sees a baseline test run across the assistants they use.
- Grounded in (2025-2026 signals)
- Agent Skills open standard at agentskills.io since December 18, 2025, with about 40 supporting products by June 2026; MCP at 10,000 active servers as of the December 9, 2025 donation; a skill-creator skill already writes skills, so the corpus of untested skills grows daily.
- What it rides
- 'Agent Skills: launched October 16, 2025, an open standard since December 18': the brief says skills can be 'built, tested, versioned, installed, audited and sold', but the testing and versioning tooling behind that sentence does not exist yet.
- Why now
- The standard is eight months old as of August 2026 and already runs in about 40 products, so one skill now executes on many models that each update on their own schedule; every frontier release silently changes skill behavior and there is no CI equivalent to catch it.
- Wedge: first customer and entry point
- Small teams of Claude Code users who share a skills repository; the entry point is a drop-in check that runs on every commit to the skills folder, priced per skill per month, built by a 1-3 person team on top of existing assistant CLIs.
- Closest real companies, as the generator saw them
- Quantstruct (yc W25) tests and improves stale product docs, not agent skills. Mastra (yc W25) is a framework for authoring your own agents; Skillvane tests skills that run inside other companies' assistants. Confident AI (yc W25) evaluates LLM applications, not installable skill packages.
- Main risk
- Anthropic folds testing and versioning into the official skill-creator toolchain and the standard's steward owns the CI layer.
Similar startups in the directory
Companies whose pitch matches most of the concept's terms (version, control, regression, tests, skills, load, registry, skill).
Robotic crews for large-scale infrastructure construction
MintyCode: The Infrastructure Layer for Trusted AI Systems
LLMs that control robots
One connector per employee. Every AI tool and skill your company allows, and nothing else.
A personal or team assistant that works while you sleep
Agentic infrastructure to power AI-native companies.
Hire Zalos AI workers to run your finance department
AI models that teach robots new skills in hours
Fluid Wire Robotics develops high-performance force-controllable robotic arms for operations in critical environments. Based on proprietary "Fluid Wire" technology, they enable advanced manipulation in oil & gas facilities, underwater areas, nuclear power plants, mines and clean-rooms.
Jobo is a platform in Ivory Cost offering rapid, reliable, and automated temporary employment services. They use AI to preselect candidates within 24 hours and ensure all temporary workers are trained before starting their roles.
Micropsi industries provides high-end machine learning solutions for robotics and process control.
Trains AI-powered physical skills for robots
Public money in this direction
US federal grants and open opportunities matched to the concept's terms.
NIH / NIA · SBIR phase I · $349K
National Science Foundation · SBIR Phase II · $1M
NIH / NIMH · SBIR phase I · $417K
NIH / NIBIB · STTR phase I · $320K
NIH / NCHHSTP · SBIR phase I · $307K
National Science Foundation · STTR Phase I · $305K
National Science Foundation · M3X - Mind, Machine, and Motor, Special Initiatives, Dynamics, Control and System D · $809K
NIH / NCI · SBIR phase II · $404K
Other concepts in this collection
- SkillproofRegression testing for the Agent Skills you actually depend on
- ProvenaryScan third-party skills and MCP servers before you let them touch your data
- LedgerkitVersioned skill packs that make a solo CPA's assistant work like a tax practice
- VendfoldLicensing, signing and auto-update infrastructure for people who sell Agent Skills
- PackroomOne shared skill library for a team where everyone runs their own agent
- TokentabPer-skill cost, routing and drift telemetry for the person who runs AI all day
- ThreadkeepA memory vault you own that every assistant you run can read
- RelayfileHand a running task from Claude Code to Codex without losing state
- MeterhouseOne budget, meter and kill switch for every agent you run
- AttestlyAudit trail and approval inbox for the agents you run at work
- CrewlineA shared board where each teammate's agents pick up each other's work
- WardkeySecurity scanner that finds and fixes exposed keys in vibe-coded apps
- StillupUptime and error monitoring that answers in fix prompts, not stack traces
- CopystoneAutomatic backups and one-click restore for apps built without engineers
- GroundskeepMonthly maintenance for shipped vibe-coded apps, applied as reviewable patches
- TillhousePayments, sales tax and refunds as one drop-in for non-developer founders
- SpendgateMeter, cap and route the AI spend inside apps vibe coders shipped
- DryloopRehearsal mode for the automations a small business owner builds alone
- MeterlyOne metered key with spend caps for every AI step you run
- FlowmedicWatches your automations, explains failures in plain English, proposes the fix
- ScrubdeckA data-cleaning step any workflow can call, with rules the owner keeps
- OpshandTurns your written SOPs into versioned Agent Skills with tests included
- CrewtraceShared visibility when five people at one business each run their own automations
- VeraciteCitation verification and AI work records for solo attorneys who draft with Claude
- TickstoneTurns a solo CPA's AI sessions into reviewable workpapers with tickmarks and source trails
- ChartproofA verification layer for physicians who use AI on clinical notes under their own license
- CoverlensPolicy-form verification for independent insurance agents who quote with AI
- MethodkitSolo consultants package their methodology as versioned Agent Skills they own and resell
- AttestrailTamper-evident logs of every AI action, built for licensed professionals' liability files
- ScrublineLocal redaction proxy that makes your personal AI accounts safe for work data
- StipendlyTurn personal Claude Max and ChatGPT Pro seats into managed employer stipends
- TollgateA policy gateway between your assistant and every MCP server it touches
- SkillvetScan, pin and approve Agent Skills before they touch company data
- DaylightSelf-serve shadow AI registry and policy for companies with no security team
- LedgerlineRightsizing dashboard for everyone paying for AI out of their own pocket
- SwitchyardOne metered endpoint with routing, fallback and per-person caps for tiny teams
- HearthmeterUsage budgets and one bill for the household that shares AI plans
- SeatcaseMeasures who on your team earns a Max seat and who wastes one
- TokencairnProfiler that shows what each installed skill and MCP server really costs
- FusegateBudget caps, fallback and kill switches for automations you run yourself
- SkillbenchRegression testing for Agent Skills before every model and skill update
- CitelockVerifies every citation in AI-drafted work before a licensed professional signs it
- MiddlegateA local gateway where you set the rules for what your MCP servers can do
- DriftwatchCatches output drift in the automations small operators wired themselves
- ShipcheckPre-launch review gates non-technical builders run on their own vibe-coded apps
- TracelineA claim-level provenance trail for every number in an AI-assisted report
- DrillyardScored practice repos where you learn to drive coding agents well
- PassrateA proctored AI operation exam scored from your real agent transcripts
- PatchcraftDebugging drills that teach non-technical builders to maintain what they vibe coded
- SkillsmithA workshop for writing, testing and versioning Agent Skills that actually hold up
- TickmarkSynthetic client caseloads where CPAs drill AI-assisted work before trying it on real clients
- PostgameAn MCP server that scores your own agent sessions and drills your weakest habits
- CitegridEvery number in your published research links to a source snapshot you verified
- MnemosYour research corpus as a private MCP server every assistant can query
- MeterlineModel routing and cost accounting for one person's AI research pipeline
- SkillcaskVersion, test, and sell your expertise as licensed Agent Skills
- StackfeedA personal data pipeline that repairs itself when sources change
- ClaimboardA shared evidence ledger for small teams where everyone runs their own agent
Fictional concept generated 2026-08-26 by claude-fable-5 from the collection's brief and MarkosWeb data. Treat it as a research prompt, not a plan.