New startup ideas · AI for people who run the AI themselves · Learning to wield AI
startup concept
Skillsmith
A workshop for writing, testing and versioning Agent Skills that actually hold up
Skillsmith is a practice environment for authoring SKILL.md skills: you write a skill, it runs your skill against a battery of held-out tasks across Claude, Codex and Cursor, and shows where the instructions fail, with graded exercises that build from a one-file skill to a versioned pack with scripts and evals.
- Software subscription
- Consumer
- Rides 'Agent Skills
7
similar startups, last 2 years (17 all-time)
yes
8 matching federal grants and programs
Direction supported by government programs and grants
Test it before you build it
$400 · 3 weeks · 25 prospects
For $400 and 3 weeks, learn whether people who ship Agent Skills will pay $40 a month to watch those skills fail on assistants they do not use
Riskiest assumption · An author who ships skills to clients or their own stack will pay $40 a month for cross-assistant defect reports instead of trusting the skill-creator built into Claude
1Focus group: who and where
A consultant or Claude Max subscriber who has published or delivered at least one SKILL.md skill - to a client, a team, or a public repo - and whose reputation takes the hit when the skill misfires somewhere they never tested
where to find 25 · r/ClaudeAI on Reddit, where skill authors post their packs and failure stories; GitHub search for repositories containing SKILL.md files with commits in the last 90 days, a ready-made list of authors with contact links; AI Tinkerers meetups in NYC, SF and Austin, where the same people demo agent setups in person
2Sell first, build later
A founding slot in the September test cohort: submit one skill per week, get a hand-run defect report across Claude, Codex and Cursor within 5 days of each submission, starting the week you pay - the battery is run manually, no product exists yet
the ask · $40 per month, first month prepaid, founding rate locked for a year against a planned $60 list price
a real yes · A real yes is $40 charged and a skill file submitted; stars on the teardown posts, 'this should exist' replies, and consultants asking for a free sample report do not count
3Small experiments
The first one attacks the riskiest assumption; each ends with a number that says whether to run the next.
1. Founding tester presale
$100 · 14 days
Offer a 15-slot founding month: send one skill, the founder runs it by hand against a 20-task held-out battery on Claude, Codex and Cursor, and returns a defect list within 5 days - where instructions get ignored, where behavior drifts between assistants. Contact 25 authors pulled from the GitHub SKILL.md search and r/ClaudeAI with the outreach script; close at $40 for the first month prepaid.
keep going if · 8 of 25 authors prepay $40 and 5 submit a skill within the first week
2. Consultant resale interviews
$0 · 10 days
Book 10 calls with authors who write skills for clients (found via the same GitHub list and r/ClaudeAI posts advertising services). Ask what a skill failure at a client has cost them, whether they would attach a Skillsmith defect report to deliverables, and what they would pay - then make the $40 offer live on the call.
keep going if · 6 of 10 name a real failure that cost them client trust, and 4 of 10 buy on the call
3. Public skill teardown
$300 · 7 days
Pick 3 widely shared public skills from GitHub, run the manual battery on all three assistants (one month of each tool plus API credits), and publish the defect writeups on r/ClaudeAI with a submission link. This tests whether visible cross-assistant failures pull authors in without outreach.
keep going if · The posts generate 10 or more skill submissions or purchase-page visits that convert to 3 paid founding slots
4Collect a deposit up front
Tesla took $1,000 refundable reservations for the Model 3 and $100 for the Cybertruck before building either: the deposit is the measurement, not the revenue.
$40
per prospect, refundable
how · First month prepaid at checkout on the cohort page - for a $40-a-month consumer subscription the first month is the deposit, and anything smaller would make the yes too cheap to mean anything set up: Stripe Checkout ↗
what it reserves · One of 15 founding slots, the $40 rate locked for 12 months, and a guaranteed 5-day turnaround on each weekly report during the founding month
refund · Refunded in full if the first defect report does not arrive within 5 days of submission or if the author asks before receiving it
target · 8 prepaid first months from 25 conversations within 21 days
Go: build it if
8 of 25 authors prepay $40, at least 5 submit a skill in week one, and at least 4 submit a second skill after their first report - they pay and they come back
Kill: stop if
3 or fewer prepay from 25 conversations, or paid users read one report and never resubmit - meaning Claude's built-in skill-creator already feels good enough, which is the stated risk confirmed
5 Scripts to run itoutreach message, landing copy, deposit terms · click to open
outreach message
You published a skill pack, and right now you only know it works in the setup where you wrote it - the first person to run it under Codex or Cursor finds the failures for you. I run Skillsmith: send me one skill, I run it against a 20-task held-out battery on Claude, Codex and Cursor, and send you a defect list within 5 days. Founding rate is $40 for the first month, 15 slots, locked for a year. Up for a 20-minute call to pick which skill goes through first?
landing page
Your skill works in your setup. Find out where it doesn't. $40 a month: weekly defect reports on your skills across Claude, Codex and Cursor, 5-day turnaround. Prepay month one and submit your first skill.
deposit terms
$40 prepays your first month and holds one of 15 founding slots at that rate, locked for 12 months. Submit one skill per week; each defect report arrives within 5 days of submission. Full refund if your first report is late or if you ask before it arrives.
Would you run this test?
One tap. The yes-share feeds the Demand pillar of this idea's score; nobody sees who answered.
Budgets are out-of-pocket estimates for a team of one to three, US market. Size the deposit to the deal, and check the terms before taking money in a regulated line.
Scorecard
Ranked against every idea in the catalog: trend, demand and 100x potential from the corpus, competition relative to the other ideas. A generated concept has no judges or swipes yet, so its pillars use the data signals only.
72
Idea Score, 0-100 · raw 44.3 x 1.61
Active
competition: more crowded than 69% of ideas · headwind x0.65
+4.9
government priorities, secondary (170 matching grants)
Trend
47
Is the wave forming now? 2025-26 entrants vs 2023-24, rounds since 2025, the sector's live-batch direction, the 2026 trend analyst.
- Entrants 2025-26 vs 2023-24 (similar companies)90
- Rounds announced 2025+ in the sector0
- Sector direction (live batch)50
Demand
65
Does anyone want it? YC's current RFS, companies already paid for something similar, the operator judge, founders' yes-rate in decks, readers who would run the test.
- YC asks for it (current RFS: idea / sector)30
- Someone already pays (similar companies, recent / all-time)100
100x potential
72
Can it return a fund? The venture judge (double weight), market-size and moat axes, neighbours still alive, the technologist judge.
- Neighbours still alive72
Score = 100 x cbrt(Trend x Demand x 100x) x (1 - 0.5 x crowding) + government bonus (max 5), calibrated so the 95th-percentile idea scores 90 (order never changes). A geometric mean: a weak pillar cannot be papered over. Percentiles are among the 382 ideas in the catalog; the terms matched were workshop, writing, testing, versioning, skills, hold, practice, environment.
The concept in full
- What
- Skillsmith is a practice environment for authoring SKILL.md skills: you write a skill, it runs your skill against a battery of held-out tasks across Claude, Codex and Cursor, and shows where the instructions fail, with graded exercises that build from a one-file skill to a versioned pack with scripts and evals. Power users and consultants who want to sell or ship skills get a pass or fail report on their first skill within the hour.
- Grounded in (2025-2026 signals)
- Agent Skills launched October 16, 2025 and became an open standard at agentskills.io on December 18, 2025, supported by about 40 products including Claude, OpenAI Codex, GitHub Copilot, VS Code, Cursor, Gemini CLI and Goose by June 2026; MCP's donation to the Agentic AI Foundation on December 9, 2025 with 10,000 active servers shows the parallel curve; the keyword 'prompt' appears in 62 companies from the last two years, but skill authoring craft is untracked.
- What it rides
- Rides 'Agent Skills: launched October 16, 2025, an open standard since December 18' directly: a new authoring discipline with about 40 supporting products by June 2026 and no place to learn it, the same position code review tools held when open source standardized software.
- Why now
- Between December 18, 2025 and June 2026 the skill format went from one vendor to about 40 products, which is the exact window where an authoring craft forms; whoever defines what a well-tested skill looks like in 2026 sets the norms before the ecosystem hardens.
- Wedge: first customer and entry point
- First customers are the consultants and $200-a-month Max subscribers already writing skills for clients or their own stack; the wedge is the cross-product test battery, run your skill against three assistants and get a defect list, which a two person team can build on the open standard and price at $40 a month.
- Closest real companies, as the generator saw them
- Confident AI evaluates LLM applications for engineering teams; Skillsmith evaluates the skill artifact and trains its author. The Prompting Company optimizes how products appear to agents, a different artifact entirely. No tracked company teaches skill authoring.
- Main risk
- The skill-creator skill inside Claude gets good enough that authors never feel the need to test or learn.
Similar startups in the directory
Companies whose pitch matches most of the concept's terms (workshop, writing, testing, versioning, skills, hold, practice, environment).
Healthcare AI
We are developing a NoCode application aimed to help SMEs in their process of aumatization. Companies who'll use our software will be able to configure, program and control a robotic cell without write a single line of code, working in a user-friendly and easy-to-understand environment.
AI Agent for Therapists
Reason Machines builds Reason Agent, the autonomous software engineer…
Xpressive AI — real-time training with AI avatars that help you practice conversations, build skills and track your improvement
Talk Labs is a platform that helps companies train customer-facing and frontline employees through realistic AI voice trainers: businesses can either use ready-to-deploy coaches or upload their own scripts, service standards, sales methodologies, rare edge cases, and internal materials, which the system then turns into interactive voice simulations where employees speak with AI as if they were talking to a real customer, guest, or buyer, practice in a safe environment, and receive instant personalized feedback on mistakes, argumentation, tone, and service quality, while managers and L&D teams get analytics on skills, gaps, progress, training quality, and its impact on the business; Talk Labs replaces one-off training sessions and poorly scalable manual practice with an always-available, highly scalable, measurable, and quickly adaptable AI infrastructure for improving sales and service in industries such as hospitality, retail, banking, and telecom.
Luna Robotics producing affordable night vision for FPV kamikaze drones.
Personalized prevention using genetics and AI
Developer of telco cloud training platform intended to provide hands-on practical learning experiences. The company platform offers testing and validation processes staff training with round-the-cloud access granted to anyone who buys a license and tokens, enabling companies to set up a telco cloud environment by plugging into on-demand labs that allow users to test various scenarios that are relevant to the technology stack.
Automated e2e screenshot testing without writing or maintaining tests
Build webapps faster with 10x better CI/CD + preview envs
The Computer Science Proficiency Assessment (CSPA™) is a…
Public money in this direction
US federal grants and open opportunities matched to the concept's terms.
NIH / NIGMS · SBIR phase II · $1M
National Science Foundation · SBIR Phase II · $1M
NIH / NIA · SBIR phase I · $505K
NIH / NINR · STTR phase I · $307K
NIH / NICHD · SBIR phase II · $739K
NIH / NIA · STTR phase I · $307K
NIH / NIBIB · STTR phase I · $350K
NIH / NHLBI · SBIR phase II · $1M
Other concepts in this collection
- SkillproofRegression testing for the Agent Skills you actually depend on
- ProvenaryScan third-party skills and MCP servers before you let them touch your data
- LedgerkitVersioned skill packs that make a solo CPA's assistant work like a tax practice
- VendfoldLicensing, signing and auto-update infrastructure for people who sell Agent Skills
- PackroomOne shared skill library for a team where everyone runs their own agent
- TokentabPer-skill cost, routing and drift telemetry for the person who runs AI all day
- ThreadkeepA memory vault you own that every assistant you run can read
- RelayfileHand a running task from Claude Code to Codex without losing state
- MeterhouseOne budget, meter and kill switch for every agent you run
- AttestlyAudit trail and approval inbox for the agents you run at work
- SkillvaneVersion control and regression tests for the skills your agents load
- CrewlineA shared board where each teammate's agents pick up each other's work
- WardkeySecurity scanner that finds and fixes exposed keys in vibe-coded apps
- StillupUptime and error monitoring that answers in fix prompts, not stack traces
- CopystoneAutomatic backups and one-click restore for apps built without engineers
- GroundskeepMonthly maintenance for shipped vibe-coded apps, applied as reviewable patches
- TillhousePayments, sales tax and refunds as one drop-in for non-developer founders
- SpendgateMeter, cap and route the AI spend inside apps vibe coders shipped
- DryloopRehearsal mode for the automations a small business owner builds alone
- MeterlyOne metered key with spend caps for every AI step you run
- FlowmedicWatches your automations, explains failures in plain English, proposes the fix
- ScrubdeckA data-cleaning step any workflow can call, with rules the owner keeps
- OpshandTurns your written SOPs into versioned Agent Skills with tests included
- CrewtraceShared visibility when five people at one business each run their own automations
- VeraciteCitation verification and AI work records for solo attorneys who draft with Claude
- TickstoneTurns a solo CPA's AI sessions into reviewable workpapers with tickmarks and source trails
- ChartproofA verification layer for physicians who use AI on clinical notes under their own license
- CoverlensPolicy-form verification for independent insurance agents who quote with AI
- MethodkitSolo consultants package their methodology as versioned Agent Skills they own and resell
- AttestrailTamper-evident logs of every AI action, built for licensed professionals' liability files
- ScrublineLocal redaction proxy that makes your personal AI accounts safe for work data
- StipendlyTurn personal Claude Max and ChatGPT Pro seats into managed employer stipends
- TollgateA policy gateway between your assistant and every MCP server it touches
- SkillvetScan, pin and approve Agent Skills before they touch company data
- DaylightSelf-serve shadow AI registry and policy for companies with no security team
- LedgerlineRightsizing dashboard for everyone paying for AI out of their own pocket
- SwitchyardOne metered endpoint with routing, fallback and per-person caps for tiny teams
- HearthmeterUsage budgets and one bill for the household that shares AI plans
- SeatcaseMeasures who on your team earns a Max seat and who wastes one
- TokencairnProfiler that shows what each installed skill and MCP server really costs
- FusegateBudget caps, fallback and kill switches for automations you run yourself
- SkillbenchRegression testing for Agent Skills before every model and skill update
- CitelockVerifies every citation in AI-drafted work before a licensed professional signs it
- MiddlegateA local gateway where you set the rules for what your MCP servers can do
- DriftwatchCatches output drift in the automations small operators wired themselves
- ShipcheckPre-launch review gates non-technical builders run on their own vibe-coded apps
- TracelineA claim-level provenance trail for every number in an AI-assisted report
- DrillyardScored practice repos where you learn to drive coding agents well
- PassrateA proctored AI operation exam scored from your real agent transcripts
- PatchcraftDebugging drills that teach non-technical builders to maintain what they vibe coded
- TickmarkSynthetic client caseloads where CPAs drill AI-assisted work before trying it on real clients
- PostgameAn MCP server that scores your own agent sessions and drills your weakest habits
- CitegridEvery number in your published research links to a source snapshot you verified
- MnemosYour research corpus as a private MCP server every assistant can query
- MeterlineModel routing and cost accounting for one person's AI research pipeline
- SkillcaskVersion, test, and sell your expertise as licensed Agent Skills
- StackfeedA personal data pipeline that repairs itself when sources change
- ClaimboardA shared evidence ledger for small teams where everyone runs their own agent
Fictional concept generated 2026-08-26 by claude-fable-5 from the collection's brief and MarkosWeb data. Treat it as a research prompt, not a plan.