New startup ideas · AI for people who run the AI themselves · Learning to wield AI
startup concept
Postgame
An MCP server that scores your own agent sessions and drills your weakest habits
Postgame plugs into whatever assistant you already run via MCP, keeps your transcripts in a store you own, and grades each session on cost, verification, delegation and rework, then generates personalized drills from your actual failure patterns, the prompt you re-ran four times, the output you accepted unverified.
- Infrastructure and APIs
- Consumer
- Rides 'MCP donated to the Agentic AI Foundation' (December 9
1
similar startups, last 2 years (2 all-time)
no
no public money matching the concept's terms
Test it before you build it
$250 · 3 weeks · 40 prospects
$250 and three weeks to prove that people paying $100-200 a month for AI capacity will prepay $49 to see their own agent sessions graded, before the MCP server exists.
Riskiest assumption · A subscriber already paying $100-200 a month for AI capacity believes their own misuse wastes enough money that they will pay $49 up front for a graded replay of their sessions, rather than expecting the assistant to critique itself for free.
1Focus group: who and where
A developer or technical solo operator on Claude Max or ChatGPT Pro who runs agent sessions daily, posts token-spend or usage-limit screenshots, and suspects rework and unverified output are eating their $1,200-2,400 a year in capacity but has no number for it.
where to find 40 · r/ClaudeAI and r/ChatGPTPro, where subscribers post usage screenshots weekly (communities); the Claude Developers Discord (channel); Show HN and the agent-workflow posters on X (channel); AI Tinkerers meetups, which run monthly chapters in SF, NYC and Seattle (event).
2Sell first, build later
A founding quarter of Postgame, delivered manually before any MCP server exists: weekly hand-built scorecards on the member's real sessions covering cost, verification, delegation and rework, plus three personalized drills generated from their own failure patterns. Starts October 15.
the ask · $49 for the first three months, converting to usage pricing afterward at a locked founding rate
a real yes · A real yes is $49 charged. Waitlist emails, Discord praise, HN upvotes and transcripts shared for the free scorecard do not count.
3Small experiments
The first one attacks the riskiest assumption; each ends with a number that says whether to run the next.
1. Hand-scored replay calls
$100 · 10 days
DM 25 people who posted usage screenshots on r/ClaudeAI, r/ChatGPTPro or the Claude Developers Discord and recruit 15 to share one week of exported session transcripts. The founder hand-builds each scorecard: re-run count, outputs accepted unverified, estimated dollars lost to rework. Deliver it on a 20-minute call and end with the $49 founding-quarter ask.
keep going if · 9 of 15 recruits actually send transcripts; 5 of 15 pay $49 on or within a day of the call
2. Show HN priced waitlist
$150 · 7 days
Publish one anonymized scorecard as a write-up: this user wasted $84 of a $200 month on re-runs, here is the grading rubric. Post it as Show HN and to r/ClaudeAI, linking a page that sells the $49 founding quarter starting October 15, capped at 25 seats.
keep going if · 300 unique visitors, 25 emails captured, 8 paying $49 (2.7% of visitors)
3. Second-week retention check
$0 · 7 days
For everyone who paid, deliver week one's scorecard, then wait: do they send week two's transcripts without prompting? A coaching product that people stop feeding after one report fails even if it sells once. The founder just tracks submissions.
keep going if · 8 of the first 10 paying users submit a second week of transcripts unprompted
4Collect a deposit up front
Tesla took $1,000 refundable reservations for the Model 3 and $100 for the Cybertruck before building either: the deposit is the measurement, not the revenue.
$49
per prospect, refundable
how · $49 founding-quarter prepay by card checkout linked from the scorecard call and the landing page. Card checkout fits because this buyer already pays $100-200 a month by card for the assistant itself; a $49 charge is below their monthly tool noise. No contract, just posted terms. set up: Stripe Checkout ↗
what it reserves · One of 25 first-cohort seats starting October 15, weekly scorecards through the quarter, and the founding usage rate locked when metered billing begins.
refund · Refund on request any time before the first scorecard is delivered.
target · 10 prepays of $49 from 40 prospects within 21 days
Go: build it if
10 prepaid founding quarters within 21 days and at least 8 of the first 10 paying users submitting a second week of transcripts unprompted: build the MCP server.
Kill: stop if
Fewer than 4 prepays from 15 calls plus 300 landing visitors, or users read their own waste number and shrug: the waste is visible but not worth $49 to fix, and the native-analytics risk wins, stop.
5 Scripts to run itoutreach message, landing copy, deposit terms · click to open
outreach message
You posted your Claude token spend last week, which is exactly who I'm looking for. I'm hand-building scorecards from real session transcripts: what you re-ran four times, what you accepted unverified, what the rework cost in dollars. I'll score one week of your sessions free and walk you through it in 20 minutes, and you'll leave with the number you're wasting each month. You share only the transcripts you choose, and they stay in your own store. Up for it this week?
landing page
Your agent sessions, graded $49 founding quarter: weekly scorecards plus drills built from your own failure patterns Take one of 25 first-cohort seats, starts October 15
deposit terms
$49 covers your first three months in the founding cohort of 25, starting October 15, and locks the founding rate when usage billing begins. Refund on request any time before your first scorecard is delivered. Your transcripts stay in a store you control; we only read the sessions you send.
Would you run this test?
One tap. The yes-share feeds the Demand pillar of this idea's score; nobody sees who answered.
Budgets are out-of-pocket estimates for a team of one to three, US market. Size the deposit to the deal, and check the terms before taking money in a regulated line.
Scorecard
Ranked against every idea in the catalog: trend, demand and 100x potential from the corpus, competition relative to the other ideas. A generated concept has no judges or swipes yet, so its pillars use the data signals only.
49
Idea Score, 0-100 (partial) · raw 30.5 x 1.61
Warm
competition: more crowded than 32% of ideas · headwind x0.84
+0.0
government priorities, secondary (0 matching grants)
Trend
41
Is the wave forming now? 2025-26 entrants vs 2023-24, rounds since 2025, the sector's live-batch direction, the 2026 trend analyst.
- Entrants 2025-26 vs 2023-24 (similar companies)74
- Rounds announced 2025+ in the sector0
- Sector direction (live batch)50
Demand
23
Does anyone want it? YC's current RFS, companies already paid for something similar, the operator judge, founders' yes-rate in decks, readers who would run the test.
- YC asks for it (current RFS: idea / sector)30
- Someone already pays (similar companies, recent / all-time)17
100x potential
50
Can it return a fund? The venture judge (double weight), market-size and moat axes, neighbours still alive, the technologist judge.
- no signal yet, taken as 50
Score = 100 x cbrt(Trend x Demand x 100x) x (1 - 0.5 x crowding) + government bonus (max 5), calibrated so the 95th-percentile idea scores 90 (order never changes). A geometric mean: a weak pillar cannot be papered over. Percentiles are among the 382 ideas in the catalog; the terms matched were mcp, server, scores, sessions, drills, weakest, plugs, whatever.
The concept in full
- What
- Postgame plugs into whatever assistant you already run via MCP, keeps your transcripts in a store you own, and grades each session on cost, verification, delegation and rework, then generates personalized drills from your actual failure patterns, the prompt you re-ran four times, the output you accepted unverified. A daily AI user installs the server, replays last week's sessions, and sees their first scorecard within the hour; pricing is by usage.
- Grounded in (2025-2026 signals)
- MCP donated December 9, 2025 with 97 million monthly SDK downloads and 10,000 active servers; Claude Max launched April 9, 2025 at $100 and $200 a month for people who 'collaborate with Claude extensively'; Woz (YC W25) cuts Claude Code token cost 50% via plugin, proving session-level waste is real and measurable; MIT NANDA (August 2025): users call the same models reliable in their own hands and unreliable in enterprise systems.
- What it rides
- Rides 'MCP donated to the Agentic AI Foundation' (December 9, 2025): with 97 million monthly SDK downloads and first-class support in ChatGPT, Claude, Cursor, Gemini, Copilot and VS Code, one server can observe a user's practice across every assistant they run, which is what makes coaching from your own data possible.
- Why now
- The $100 and $200 Max tiers that opened April 9, 2025 created users spending $1,200 to $2,400 a year on capacity with zero feedback on how well they use it, and MCP's December 9, 2025 standardization means one integration now reaches every major assistant instead of six.
- Wedge: first customer and entry point
- First customers are Max and ChatGPT Pro subscribers who already track their token spend; the entry point is a free week of scored replays showing dollars wasted on rework, then usage pricing that stays a fraction of the waste it removes, buildable by one person as a single MCP server.
- Closest real companies, as the generator saw them
- Woz reduces token cost automatically inside Claude Code but teaches nothing and covers one tool; Confident AI observes LLM apps for engineering teams, not individuals. jo is a personal agent that does work rather than coaching how you work. None score the human across assistants.
- Main risk
- Assistants start shipping native usage analytics and self-critique, collapsing the space between the model and the coach.
Similar startups in the directory
Companies whose pitch matches most of the concept's terms (mcp, server, scores, sessions, drills, weakest, plugs, whatever).
Public money in this direction
US federal grants and open opportunities matched to the concept's terms.
No matching grants or programs tracked; 19 startup-relevant grants exist in Horizontal AI assistants overall.
Other concepts in this collection
- SkillproofRegression testing for the Agent Skills you actually depend on
- ProvenaryScan third-party skills and MCP servers before you let them touch your data
- LedgerkitVersioned skill packs that make a solo CPA's assistant work like a tax practice
- VendfoldLicensing, signing and auto-update infrastructure for people who sell Agent Skills
- PackroomOne shared skill library for a team where everyone runs their own agent
- TokentabPer-skill cost, routing and drift telemetry for the person who runs AI all day
- ThreadkeepA memory vault you own that every assistant you run can read
- RelayfileHand a running task from Claude Code to Codex without losing state
- MeterhouseOne budget, meter and kill switch for every agent you run
- AttestlyAudit trail and approval inbox for the agents you run at work
- SkillvaneVersion control and regression tests for the skills your agents load
- CrewlineA shared board where each teammate's agents pick up each other's work
- WardkeySecurity scanner that finds and fixes exposed keys in vibe-coded apps
- StillupUptime and error monitoring that answers in fix prompts, not stack traces
- CopystoneAutomatic backups and one-click restore for apps built without engineers
- GroundskeepMonthly maintenance for shipped vibe-coded apps, applied as reviewable patches
- TillhousePayments, sales tax and refunds as one drop-in for non-developer founders
- SpendgateMeter, cap and route the AI spend inside apps vibe coders shipped
- DryloopRehearsal mode for the automations a small business owner builds alone
- MeterlyOne metered key with spend caps for every AI step you run
- FlowmedicWatches your automations, explains failures in plain English, proposes the fix
- ScrubdeckA data-cleaning step any workflow can call, with rules the owner keeps
- OpshandTurns your written SOPs into versioned Agent Skills with tests included
- CrewtraceShared visibility when five people at one business each run their own automations
- VeraciteCitation verification and AI work records for solo attorneys who draft with Claude
- TickstoneTurns a solo CPA's AI sessions into reviewable workpapers with tickmarks and source trails
- ChartproofA verification layer for physicians who use AI on clinical notes under their own license
- CoverlensPolicy-form verification for independent insurance agents who quote with AI
- MethodkitSolo consultants package their methodology as versioned Agent Skills they own and resell
- AttestrailTamper-evident logs of every AI action, built for licensed professionals' liability files
- ScrublineLocal redaction proxy that makes your personal AI accounts safe for work data
- StipendlyTurn personal Claude Max and ChatGPT Pro seats into managed employer stipends
- TollgateA policy gateway between your assistant and every MCP server it touches
- SkillvetScan, pin and approve Agent Skills before they touch company data
- DaylightSelf-serve shadow AI registry and policy for companies with no security team
- LedgerlineRightsizing dashboard for everyone paying for AI out of their own pocket
- SwitchyardOne metered endpoint with routing, fallback and per-person caps for tiny teams
- HearthmeterUsage budgets and one bill for the household that shares AI plans
- SeatcaseMeasures who on your team earns a Max seat and who wastes one
- TokencairnProfiler that shows what each installed skill and MCP server really costs
- FusegateBudget caps, fallback and kill switches for automations you run yourself
- SkillbenchRegression testing for Agent Skills before every model and skill update
- CitelockVerifies every citation in AI-drafted work before a licensed professional signs it
- MiddlegateA local gateway where you set the rules for what your MCP servers can do
- DriftwatchCatches output drift in the automations small operators wired themselves
- ShipcheckPre-launch review gates non-technical builders run on their own vibe-coded apps
- TracelineA claim-level provenance trail for every number in an AI-assisted report
- DrillyardScored practice repos where you learn to drive coding agents well
- PassrateA proctored AI operation exam scored from your real agent transcripts
- PatchcraftDebugging drills that teach non-technical builders to maintain what they vibe coded
- SkillsmithA workshop for writing, testing and versioning Agent Skills that actually hold up
- TickmarkSynthetic client caseloads where CPAs drill AI-assisted work before trying it on real clients
- CitegridEvery number in your published research links to a source snapshot you verified
- MnemosYour research corpus as a private MCP server every assistant can query
- MeterlineModel routing and cost accounting for one person's AI research pipeline
- SkillcaskVersion, test, and sell your expertise as licensed Agent Skills
- StackfeedA personal data pipeline that repairs itself when sources change
- ClaimboardA shared evidence ledger for small teams where everyone runs their own agent
Fictional concept generated 2026-08-26 by claude-fable-5 from the collection's brief and MarkosWeb data. Treat it as a research prompt, not a plan.