New startup ideas · AI for people who run the AI themselves · Learning to wield AI
startup concept
Drillyard
Scored practice repos where you learn to drive coding agents well
Drillyard is a bank of sandboxed, deliberately broken codebases with timed drills: fix the bug, ship the feature, refactor safely, but you must do it by directing Cursor or Claude Code, not by hand.
- Software subscription
- Consumer
- Rides 'Cursor
11
similar startups, last 2 years (37 all-time)
yes
8 matching federal grants and programs
Direction supported by government programs and grants
Test it before you build it
$200 · 3 weeks · 30 prospects
For $200 and 3 weeks, prove that 10 of 30 developers already paying for Cursor or Claude Code will prepay $39 for a scored drill league instead of treating daily work as free practice.
Riskiest assumption · Developers already paying $20-200 a month for Cursor or Claude Code believe scored practice will make them measurably better and will pay for it, rather than treating their daily work as all the practice they need
1Focus group: who and where
An individual developer or technical solo builder in the US who pays $20-200 a month for Cursor or a Claude Max plan, posts their agent workflows publicly, and suspects they are using a tenth of what they pay for
where to find 30 · The r/cursor and r/ClaudeAI subreddits (communities where these operators compare workflows), the official Cursor forum at forum.cursor.com (the channel where paying users ask how to get more out of the tool), and AI Tinkerers meetups in NYC, Seattle and San Francisco (in-person events this exact cohort attends)
2Sell first, build later
A seat in the 25-person founding league starting September 21: 4 timed drills over 2 weeks on deliberately broken repos, each run hand-scored on outcome, token cost and verification, with written feedback on every run
the ask · $39 prepaid per league seat; founding subscribers lock $15 per month for the first year after the league
a real yes · A real yes is $39 charged before the league starts; free-drill completions, leaderboard screenshots shared on X and 'ship this' comments do not count
3Small experiments
The first one attacks the riskiest assumption; each ends with a number that says whether to run the next.
1. Prepaid drill league pre-sale
$100 · 12 days
Pre-sell a 2-week founding league starting September 21, 2026: 4 timed drills on deliberately broken repos the founders prepare by hand, driven agent-only, hand-scored on outcome, token spend and whether the operator verified before accepting. Pitch via one post each in r/cursor and the Cursor forum plus direct messages to 30 people who recently posted their agent workflows. $39 prepaid, 25 seats.
keep going if · 10 prepaid seats within 12 days of the first post
2. Free public drill funnel
$0 · 7 days
Publish one free drill as a public GitHub repo with a broken codebase, a timed task, and a submission form; hand-score every run and email back a shareable score card. Post it as a Show HN and in r/ClaudeAI the same week, with the paid league offer on the score card.
keep going if · 100 drill starts, 30% submit a finished run, and 15% of finishers click through to the league offer
3. Score-credibility interviews
$100 · 5 days
Book 10 calls with free-drill finishers who did not buy a seat, at a $10 gift card each. Ask whether the score reflected their actual skill, what would make them practice weekly, and whether daily work already feels like enough practice - the belief that kills the idea.
keep going if · 6 of 10 say the score changed how they would drive the agent on their next real task
4Collect a deposit up front
Tesla took $1,000 refundable reservations for the Model 3 and $100 for the Cybertruck before building either: the deposit is the measurement, not the revenue.
$39
per prospect, refundable
how · Prepaid seat through a card checkout linked from the score card and the forum posts - this buyer already pays Cursor or Anthropic monthly by card, so a $39 card checkout is the native rail and any heavier instrument would cost more trust than it buys set up: Stripe Checkout ↗
what it reserves · One of 25 seats in the September 21 league, hand-scored feedback on all 4 runs, and the founding price of $15 per month locked for a year
refund · Full refund any time before the first drill on September 21, and on request before drill 2 after that
target · 10 prepaid seats from 30 direct pitches plus two community posts, within 12 days
Go: build it if
10 or more prepaid seats, 100 or more free-drill starts with a 30% finish rate, and 6 of 10 interviewees saying the score changed how they drive their agent
Kill: stop if
Fewer than 4 seats sold and a free-drill finish rate under 25% - operators are curious but will not pay to be measured, and daily work is the only practice they want
5 Scripts to run itoutreach message, landing copy, deposit terms · click to open
outreach message
You're paying for Cursor or Claude Code and, like most of us, probably using a tenth of what it can do. I'm running a 2-week drill league starting September 21: 4 timed drills on deliberately broken repos, driven agent-only, hand-scored on outcome, token spend and whether you verified before accepting. $39, 25 seats, written feedback on every run. Want a 20-minute call to see drill 1 first, or just grab a seat?
landing page
Find out if you can actually drive a coding agent $39 buys a seat in the 2-week founding league: 4 scored drills, feedback on every run, then $15/month locked for a year Take the free drill now, or reserve one of 25 league seats
deposit terms
$39 today reserves your seat in the founding league starting September 21: 4 timed drills over 2 weeks, each hand-scored with written feedback, plus the founding price of $15/month locked for a year. Full refund any time before the first drill, and on request before drill 2.
Would you run this test?
One tap. The yes-share feeds the Demand pillar of this idea's score; nobody sees who answered.
Budgets are out-of-pocket estimates for a team of one to three, US market. Size the deposit to the deal, and check the terms before taking money in a regulated line.
Scorecard
Ranked against every idea in the catalog: trend, demand and 100x potential from the corpus, competition relative to the other ideas. A generated concept has no judges or swipes yet, so its pillars use the data signals only.
48
Idea Score, 0-100 · raw 29.7 x 1.61
Crowded
competition: more crowded than 80% of ideas · headwind x0.60
+4.5
government priorities, secondary (78 matching grants)
Trend
49
Is the wave forming now? 2025-26 entrants vs 2023-24, rounds since 2025, the sector's live-batch direction, the 2026 trend analyst.
- Entrants 2025-26 vs 2023-24 (similar companies)96
- Rounds announced 2025+ in the sector0
- Sector direction (live batch)50
Demand
65
Does anyone want it? YC's current RFS, companies already paid for something similar, the operator judge, founders' yes-rate in decks, readers who would run the test.
- YC asks for it (current RFS: idea / sector)30
- Someone already pays (similar companies, recent / all-time)100
100x potential
23
Can it return a fund? The venture judge (double weight), market-size and moat axes, neighbours still alive, the technologist judge.
- Neighbours still alive23
Score = 100 x cbrt(Trend x Demand x 100x) x (1 - 0.5 x crowding) + government bonus (max 5), calibrated so the 95th-percentile idea scores 90 (order never changes). A geometric mean: a weak pillar cannot be papered over. Percentiles are among the 382 ideas in the catalog; the terms matched were scored, practice, repos, learn, drive, coding, bank, sandboxed.
The concept in full
- What
- Drillyard is a bank of sandboxed, deliberately broken codebases with timed drills: fix the bug, ship the feature, refactor safely, but you must do it by directing Cursor or Claude Code, not by hand. Every run is scored on outcome, token cost and whether you verified the agent's work before accepting it. Solo builders and the 12% who use AI daily at work sign up, connect their own agent, and finish their first scored drill in under an hour.
- Grounded in (2025-2026 signals)
- Cursor went from $100 million annualized in January 2025 to about $1 billion in November 2025 and about $4 billion by May 2026; Agent Skills became an open standard on December 18, 2025 with about 40 products supporting it by June 2026; YC's Fall 2026 RFS includes 'The Primer'; MangoDesk (YC S25) and Halluminate (YC S25) sell RL environments, proving graded environments are buildable, but both train models, not people.
- What it rides
- Rides 'Cursor: $100 million to about $4 billion annualized in seventeen months' as the skill gap it created: millions of seats sold to people nobody ever trained, and 'Agent Skills: launched October 16, 2025, an open standard since December 18' as the drill format, since each drill ships as a SKILL.md the user's own assistant loads.
- Why now
- Cursor added roughly $3.9 billion of annualized revenue in seventeen months to May 2026, meaning a huge cohort of paying agent operators exists with no measured way to get better, while Gallup Q4 2025 shows only 12% of US workers use AI daily, so skill, not access, is the bottleneck.
- Wedge: first customer and entry point
- First customer is the individual developer already paying $20 to $200 a month for Cursor or Claude Max who suspects they use 10% of it; the entry point is one free public drill with a shareable score, which makes the leaderboard the distribution channel a two person team can ride.
- Closest real companies, as the generator saw them
- MangoDesk and Halluminate build long-horizon RL environments to train and evaluate models; Drillyard uses the same environment mechanics to train and score the human operator. Woz optimizes Claude Code token cost automatically; Drillyard teaches the user to do it themselves.
- Main risk
- Cursor or Anthropic ships built-in onboarding drills and the standalone practice range becomes a feature.
Similar startups in the directory
Companies whose pitch matches most of the concept's terms (scored, practice, repos, learn, drive, coding, bank, sandboxed).
Buyable is an auto lending platform that drives leads and co-underwrites loans allowing banks to lend to credit diverse borrowers at scale with minimal risk.
building the open agent runtime
zypl.ai applies 'no data' AI at scale to advance financial inclusion in emerging and frontier markets. Zypl.ai proprietary AI SaaS generates synthetic data to optimize credit decision models (both retail + SME) for financial institutions;
Veris is a sandbox platform that lets enterprises train and validate autonomous agents in realistic, high-fidelity simulations before deployment.
Behavioral intelligence infrastructure for anti-fraud
Create & play AI games with friends
Frontier AI Defenses for Social Engineering Attacks
Cobi is the AI-powered intelligence layer for financial institutions, transforming every customer interaction into a personalised, revenue-driving opportunity.
AI that Reclaims Physician Time for Patients
Issue and verify ID Credentials with AI & Quantum Security and Privacy, 25x cheaper and 50x safer than biometric verification, using Next-Generation Digital ID Wallets.
Savvi AI is a scalable, low-code platform enabling Financial Services to accelerate deployment of their own data-powered AI applications and decision agents, driving immediate ROI by reducing risk, improving operations, and growing revenue.
LendAPI enables rapid deployment of customizable embedded finance products.
Public money in this direction
US federal grants and open opportunities matched to the concept's terms.
National Science Foundation · TIP-CHIPS KTA-3 Quantum · $50K
NIH / NIMH · STTR phase II · $450K
National Science Foundation · I-Corps · $50K
National Science Foundation · I-Corps · $50K
National Science Foundation · SBIR Phase II · $1M
National Science Foundation · Software & Hardware Foundation · $530K
National Science Foundation · I-Corps · $50K
NIH / NIDCD · STTR phase I · $313K
Other concepts in this collection
- SkillproofRegression testing for the Agent Skills you actually depend on
- ProvenaryScan third-party skills and MCP servers before you let them touch your data
- LedgerkitVersioned skill packs that make a solo CPA's assistant work like a tax practice
- VendfoldLicensing, signing and auto-update infrastructure for people who sell Agent Skills
- PackroomOne shared skill library for a team where everyone runs their own agent
- TokentabPer-skill cost, routing and drift telemetry for the person who runs AI all day
- ThreadkeepA memory vault you own that every assistant you run can read
- RelayfileHand a running task from Claude Code to Codex without losing state
- MeterhouseOne budget, meter and kill switch for every agent you run
- AttestlyAudit trail and approval inbox for the agents you run at work
- SkillvaneVersion control and regression tests for the skills your agents load
- CrewlineA shared board where each teammate's agents pick up each other's work
- WardkeySecurity scanner that finds and fixes exposed keys in vibe-coded apps
- StillupUptime and error monitoring that answers in fix prompts, not stack traces
- CopystoneAutomatic backups and one-click restore for apps built without engineers
- GroundskeepMonthly maintenance for shipped vibe-coded apps, applied as reviewable patches
- TillhousePayments, sales tax and refunds as one drop-in for non-developer founders
- SpendgateMeter, cap and route the AI spend inside apps vibe coders shipped
- DryloopRehearsal mode for the automations a small business owner builds alone
- MeterlyOne metered key with spend caps for every AI step you run
- FlowmedicWatches your automations, explains failures in plain English, proposes the fix
- ScrubdeckA data-cleaning step any workflow can call, with rules the owner keeps
- OpshandTurns your written SOPs into versioned Agent Skills with tests included
- CrewtraceShared visibility when five people at one business each run their own automations
- VeraciteCitation verification and AI work records for solo attorneys who draft with Claude
- TickstoneTurns a solo CPA's AI sessions into reviewable workpapers with tickmarks and source trails
- ChartproofA verification layer for physicians who use AI on clinical notes under their own license
- CoverlensPolicy-form verification for independent insurance agents who quote with AI
- MethodkitSolo consultants package their methodology as versioned Agent Skills they own and resell
- AttestrailTamper-evident logs of every AI action, built for licensed professionals' liability files
- ScrublineLocal redaction proxy that makes your personal AI accounts safe for work data
- StipendlyTurn personal Claude Max and ChatGPT Pro seats into managed employer stipends
- TollgateA policy gateway between your assistant and every MCP server it touches
- SkillvetScan, pin and approve Agent Skills before they touch company data
- DaylightSelf-serve shadow AI registry and policy for companies with no security team
- LedgerlineRightsizing dashboard for everyone paying for AI out of their own pocket
- SwitchyardOne metered endpoint with routing, fallback and per-person caps for tiny teams
- HearthmeterUsage budgets and one bill for the household that shares AI plans
- SeatcaseMeasures who on your team earns a Max seat and who wastes one
- TokencairnProfiler that shows what each installed skill and MCP server really costs
- FusegateBudget caps, fallback and kill switches for automations you run yourself
- SkillbenchRegression testing for Agent Skills before every model and skill update
- CitelockVerifies every citation in AI-drafted work before a licensed professional signs it
- MiddlegateA local gateway where you set the rules for what your MCP servers can do
- DriftwatchCatches output drift in the automations small operators wired themselves
- ShipcheckPre-launch review gates non-technical builders run on their own vibe-coded apps
- TracelineA claim-level provenance trail for every number in an AI-assisted report
- PassrateA proctored AI operation exam scored from your real agent transcripts
- PatchcraftDebugging drills that teach non-technical builders to maintain what they vibe coded
- SkillsmithA workshop for writing, testing and versioning Agent Skills that actually hold up
- TickmarkSynthetic client caseloads where CPAs drill AI-assisted work before trying it on real clients
- PostgameAn MCP server that scores your own agent sessions and drills your weakest habits
- CitegridEvery number in your published research links to a source snapshot you verified
- MnemosYour research corpus as a private MCP server every assistant can query
- MeterlineModel routing and cost accounting for one person's AI research pipeline
- SkillcaskVersion, test, and sell your expertise as licensed Agent Skills
- StackfeedA personal data pipeline that repairs itself when sources change
- ClaimboardA shared evidence ledger for small teams where everyone runs their own agent
Fictional concept generated 2026-08-26 by claude-fable-5 from the collection's brief and MarkosWeb data. Treat it as a research prompt, not a plan.