New startup ideas · AI for people who run the AI themselves · Learning to wield AI

startup concept

Drillyard

Scored practice repos where you learn to drive coding agents well

Drillyard is a bank of sandboxed, deliberately broken codebases with timed drills: fix the bug, ship the feature, refactor safely, but you must do it by directing Cursor or Claude Code, not by hand.

11

similar startups, last 2 years (37 all-time)

yes

8 matching federal grants and programs

Direction supported by government programs and grants

Test it before you build it

$200 · 3 weeks · 30 prospects

For $200 and 3 weeks, prove that 10 of 30 developers already paying for Cursor or Claude Code will prepay $39 for a scored drill league instead of treating daily work as free practice.

Riskiest assumption · Developers already paying $20-200 a month for Cursor or Claude Code believe scored practice will make them measurably better and will pay for it, rather than treating their daily work as all the practice they need

1Focus group: who and where

An individual developer or technical solo builder in the US who pays $20-200 a month for Cursor or a Claude Max plan, posts their agent workflows publicly, and suspects they are using a tenth of what they pay for

where to find 30 · The r/cursor and r/ClaudeAI subreddits (communities where these operators compare workflows), the official Cursor forum at forum.cursor.com (the channel where paying users ask how to get more out of the tool), and AI Tinkerers meetups in NYC, Seattle and San Francisco (in-person events this exact cohort attends)

2Sell first, build later

A seat in the 25-person founding league starting September 21: 4 timed drills over 2 weeks on deliberately broken repos, each run hand-scored on outcome, token cost and verification, with written feedback on every run

the ask · $39 prepaid per league seat; founding subscribers lock $15 per month for the first year after the league

a real yes · A real yes is $39 charged before the league starts; free-drill completions, leaderboard screenshots shared on X and 'ship this' comments do not count

3Small experiments

The first one attacks the riskiest assumption; each ends with a number that says whether to run the next.

  1. 1. Prepaid drill league pre-sale

    $100 · 12 days

    Pre-sell a 2-week founding league starting September 21, 2026: 4 timed drills on deliberately broken repos the founders prepare by hand, driven agent-only, hand-scored on outcome, token spend and whether the operator verified before accepting. Pitch via one post each in r/cursor and the Cursor forum plus direct messages to 30 people who recently posted their agent workflows. $39 prepaid, 25 seats.

    keep going if · 10 prepaid seats within 12 days of the first post

  2. 2. Free public drill funnel

    $0 · 7 days

    Publish one free drill as a public GitHub repo with a broken codebase, a timed task, and a submission form; hand-score every run and email back a shareable score card. Post it as a Show HN and in r/ClaudeAI the same week, with the paid league offer on the score card.

    keep going if · 100 drill starts, 30% submit a finished run, and 15% of finishers click through to the league offer

  3. 3. Score-credibility interviews

    $100 · 5 days

    Book 10 calls with free-drill finishers who did not buy a seat, at a $10 gift card each. Ask whether the score reflected their actual skill, what would make them practice weekly, and whether daily work already feels like enough practice - the belief that kills the idea.

    keep going if · 6 of 10 say the score changed how they would drive the agent on their next real task

4Collect a deposit up front

Tesla took $1,000 refundable reservations for the Model 3 and $100 for the Cybertruck before building either: the deposit is the measurement, not the revenue.

$39

per prospect, refundable

how · Prepaid seat through a card checkout linked from the score card and the forum posts - this buyer already pays Cursor or Anthropic monthly by card, so a $39 card checkout is the native rail and any heavier instrument would cost more trust than it buys set up: Stripe Checkout

what it reserves · One of 25 seats in the September 21 league, hand-scored feedback on all 4 runs, and the founding price of $15 per month locked for a year

refund · Full refund any time before the first drill on September 21, and on request before drill 2 after that

target · 10 prepaid seats from 30 direct pitches plus two community posts, within 12 days

Go: build it if

10 or more prepaid seats, 100 or more free-drill starts with a 30% finish rate, and 6 of 10 interviewees saying the score changed how they drive their agent

Kill: stop if

Fewer than 4 seats sold and a free-drill finish rate under 25% - operators are curious but will not pay to be measured, and daily work is the only practice they want

5 Scripts to run itoutreach message, landing copy, deposit terms · click to open

outreach message

You're paying for Cursor or Claude Code and, like most of us, probably using a tenth of what it can do. I'm running a 2-week drill league starting September 21: 4 timed drills on deliberately broken repos, driven agent-only, hand-scored on outcome, token spend and whether you verified before accepting. $39, 25 seats, written feedback on every run. Want a 20-minute call to see drill 1 first, or just grab a seat?

landing page

Find out if you can actually drive a coding agent $39 buys a seat in the 2-week founding league: 4 scored drills, feedback on every run, then $15/month locked for a year Take the free drill now, or reserve one of 25 league seats

deposit terms

$39 today reserves your seat in the founding league starting September 21: 4 timed drills over 2 weeks, each hand-scored with written feedback, plus the founding price of $15/month locked for a year. Full refund any time before the first drill, and on request before drill 2.

Would you run this test?

One tap. The yes-share feeds the Demand pillar of this idea's score; nobody sees who answered.

Budgets are out-of-pocket estimates for a team of one to three, US market. Size the deposit to the deal, and check the terms before taking money in a regulated line.

Scorecard

Ranked against every idea in the catalog: trend, demand and 100x potential from the corpus, competition relative to the other ideas. A generated concept has no judges or swipes yet, so its pillars use the data signals only.

48

Idea Score, 0-100 · raw 29.7 x 1.61

Crowded

competition: more crowded than 80% of ideas · headwind x0.60

+4.5

government priorities, secondary (78 matching grants)

Trend

49

Is the wave forming now? 2025-26 entrants vs 2023-24, rounds since 2025, the sector's live-batch direction, the 2026 trend analyst.

  • Entrants 2025-26 vs 2023-24 (similar companies)96
  • Rounds announced 2025+ in the sector0
  • Sector direction (live batch)50

Demand

65

Does anyone want it? YC's current RFS, companies already paid for something similar, the operator judge, founders' yes-rate in decks, readers who would run the test.

  • YC asks for it (current RFS: idea / sector)30
  • Someone already pays (similar companies, recent / all-time)100

100x potential

23

Can it return a fund? The venture judge (double weight), market-size and moat axes, neighbours still alive, the technologist judge.

  • Neighbours still alive23

Score = 100 x cbrt(Trend x Demand x 100x) x (1 - 0.5 x crowding) + government bonus (max 5), calibrated so the 95th-percentile idea scores 90 (order never changes). A geometric mean: a weak pillar cannot be papered over. Percentiles are among the 382 ideas in the catalog; the terms matched were scored, practice, repos, learn, drive, coding, bank, sandboxed.

The concept in full

What
Drillyard is a bank of sandboxed, deliberately broken codebases with timed drills: fix the bug, ship the feature, refactor safely, but you must do it by directing Cursor or Claude Code, not by hand. Every run is scored on outcome, token cost and whether you verified the agent's work before accepting it. Solo builders and the 12% who use AI daily at work sign up, connect their own agent, and finish their first scored drill in under an hour.
Grounded in (2025-2026 signals)
Cursor went from $100 million annualized in January 2025 to about $1 billion in November 2025 and about $4 billion by May 2026; Agent Skills became an open standard on December 18, 2025 with about 40 products supporting it by June 2026; YC's Fall 2026 RFS includes 'The Primer'; MangoDesk (YC S25) and Halluminate (YC S25) sell RL environments, proving graded environments are buildable, but both train models, not people.
What it rides
Rides 'Cursor: $100 million to about $4 billion annualized in seventeen months' as the skill gap it created: millions of seats sold to people nobody ever trained, and 'Agent Skills: launched October 16, 2025, an open standard since December 18' as the drill format, since each drill ships as a SKILL.md the user's own assistant loads.
Why now
Cursor added roughly $3.9 billion of annualized revenue in seventeen months to May 2026, meaning a huge cohort of paying agent operators exists with no measured way to get better, while Gallup Q4 2025 shows only 12% of US workers use AI daily, so skill, not access, is the bottleneck.
Wedge: first customer and entry point
First customer is the individual developer already paying $20 to $200 a month for Cursor or Claude Max who suspects they use 10% of it; the entry point is one free public drill with a shareable score, which makes the leaderboard the distribution channel a two person team can ride.
Closest real companies, as the generator saw them
MangoDesk and Halluminate build long-horizon RL environments to train and evaluate models; Drillyard uses the same environment mechanics to train and score the human operator. Woz optimizes Claude Code token cost automatically; Drillyard teaches the user to do it themselves.
Main risk
Cursor or Anthropic ships built-in onboarding drills and the standalone practice range becomes a feature.

Similar startups in the directory

Companies whose pitch matches most of the concept's terms (scored, practice, repos, learn, drive, coding, bank, sandboxed).

  • Buyableplugandplay · Fintechunchecked

    Buyable is an auto lending platform that drives leads and co-underwrites loans allowing banks to lend to credit diverse borrowers at scale with minimal risk.

  • herdryc F26 · 2026 · Agent infrastructurealive

    building the open agent runtime

  • zypl.aiplugandplay PnP 2026 · 2026 · Fintechalive

    zypl.ai applies 'no data' AI at scale to advance financial inclusion in emerging and frontier markets. Zypl.ai proprietary AI SaaS generates synthetic data to optimize credit decision models (both retail + SME) for financial institutions;

  • Verisplugandplay PnP 2026 · 2026 · Agent infrastructurealive

    Veris is a sandbox platform that lets enterprises train and validate autonomous agents in realistic, high-fidelity simulations before deployment.

  • Incandoryc X26 · 2026 · Security and compliancealive

    Behavioral intelligence infrastructure for anti-fraud

  • Popsyc X26 · 2026 · Consumeralive

    Create & play AI games with friends

  • BeeSafe AIyc W26 · 2026 · Security and compliancealive

    Frontier AI Defenses for Social Engineering Attacks

  • Cobiplugandplay PnP 2025 · 2025 · B2B SaaSalive

    Cobi is the AI-powered intelligence layer for financial institutions, transforming every customer interaction into a personalised, revenue-driving opportunity.

  • Arithmedicsalchemist Alchemist Class 39 · 2025 · Vertical AI agentsunchecked

    AI that Reclaims Physician Time for Patients

  • Blerifyplugandplay PnP 2025 · 2025 · Security and compliancealive

    Issue and verify ID Credentials with AI & Quantum Security and Privacy, 25x cheaper and 50x safer than biometric verification, using Next-Generation Digital ID Wallets.

  • SAVVI AIplugandplay PnP 2025 · 2025 · AI infra and computealive

    Savvi AI is a scalable, low-code platform enabling Financial Services to accelerate deployment of their own data-powered AI applications and decision agents, driving immediate ROI by reducing risk, improving operations, and growing revenue.

  • LendAPIplugandplay PnP 2024 · 2024 · Fintechalive

    LendAPI enables rapid deployment of customizable embedded finance products.

Public money in this direction

US federal grants and open opportunities matched to the concept's terms.

Other concepts in this collection

Fictional concept generated 2026-08-26 by claude-fable-5 from the collection's brief and MarkosWeb data. Treat it as a research prompt, not a plan.