New startup ideas · AI for people who run the AI themselves · Learning to wield AI

startup concept

Postgame

An MCP server that scores your own agent sessions and drills your weakest habits

Postgame plugs into whatever assistant you already run via MCP, keeps your transcripts in a store you own, and grades each session on cost, verification, delegation and rework, then generates personalized drills from your actual failure patterns, the prompt you re-ran four times, the output you accepted unverified.

1

similar startups, last 2 years (2 all-time)

no

no public money matching the concept's terms

Test it before you build it

$250 · 3 weeks · 40 prospects

$250 and three weeks to prove that people paying $100-200 a month for AI capacity will prepay $49 to see their own agent sessions graded, before the MCP server exists.

Riskiest assumption · A subscriber already paying $100-200 a month for AI capacity believes their own misuse wastes enough money that they will pay $49 up front for a graded replay of their sessions, rather than expecting the assistant to critique itself for free.

1Focus group: who and where

A developer or technical solo operator on Claude Max or ChatGPT Pro who runs agent sessions daily, posts token-spend or usage-limit screenshots, and suspects rework and unverified output are eating their $1,200-2,400 a year in capacity but has no number for it.

where to find 40 · r/ClaudeAI and r/ChatGPTPro, where subscribers post usage screenshots weekly (communities); the Claude Developers Discord (channel); Show HN and the agent-workflow posters on X (channel); AI Tinkerers meetups, which run monthly chapters in SF, NYC and Seattle (event).

2Sell first, build later

A founding quarter of Postgame, delivered manually before any MCP server exists: weekly hand-built scorecards on the member's real sessions covering cost, verification, delegation and rework, plus three personalized drills generated from their own failure patterns. Starts October 15.

the ask · $49 for the first three months, converting to usage pricing afterward at a locked founding rate

a real yes · A real yes is $49 charged. Waitlist emails, Discord praise, HN upvotes and transcripts shared for the free scorecard do not count.

3Small experiments

The first one attacks the riskiest assumption; each ends with a number that says whether to run the next.

  1. 1. Hand-scored replay calls

    $100 · 10 days

    DM 25 people who posted usage screenshots on r/ClaudeAI, r/ChatGPTPro or the Claude Developers Discord and recruit 15 to share one week of exported session transcripts. The founder hand-builds each scorecard: re-run count, outputs accepted unverified, estimated dollars lost to rework. Deliver it on a 20-minute call and end with the $49 founding-quarter ask.

    keep going if · 9 of 15 recruits actually send transcripts; 5 of 15 pay $49 on or within a day of the call

  2. 2. Show HN priced waitlist

    $150 · 7 days

    Publish one anonymized scorecard as a write-up: this user wasted $84 of a $200 month on re-runs, here is the grading rubric. Post it as Show HN and to r/ClaudeAI, linking a page that sells the $49 founding quarter starting October 15, capped at 25 seats.

    keep going if · 300 unique visitors, 25 emails captured, 8 paying $49 (2.7% of visitors)

  3. 3. Second-week retention check

    $0 · 7 days

    For everyone who paid, deliver week one's scorecard, then wait: do they send week two's transcripts without prompting? A coaching product that people stop feeding after one report fails even if it sells once. The founder just tracks submissions.

    keep going if · 8 of the first 10 paying users submit a second week of transcripts unprompted

4Collect a deposit up front

Tesla took $1,000 refundable reservations for the Model 3 and $100 for the Cybertruck before building either: the deposit is the measurement, not the revenue.

$49

per prospect, refundable

how · $49 founding-quarter prepay by card checkout linked from the scorecard call and the landing page. Card checkout fits because this buyer already pays $100-200 a month by card for the assistant itself; a $49 charge is below their monthly tool noise. No contract, just posted terms. set up: Stripe Checkout

what it reserves · One of 25 first-cohort seats starting October 15, weekly scorecards through the quarter, and the founding usage rate locked when metered billing begins.

refund · Refund on request any time before the first scorecard is delivered.

target · 10 prepays of $49 from 40 prospects within 21 days

Go: build it if

10 prepaid founding quarters within 21 days and at least 8 of the first 10 paying users submitting a second week of transcripts unprompted: build the MCP server.

Kill: stop if

Fewer than 4 prepays from 15 calls plus 300 landing visitors, or users read their own waste number and shrug: the waste is visible but not worth $49 to fix, and the native-analytics risk wins, stop.

5 Scripts to run itoutreach message, landing copy, deposit terms · click to open

outreach message

You posted your Claude token spend last week, which is exactly who I'm looking for. I'm hand-building scorecards from real session transcripts: what you re-ran four times, what you accepted unverified, what the rework cost in dollars. I'll score one week of your sessions free and walk you through it in 20 minutes, and you'll leave with the number you're wasting each month. You share only the transcripts you choose, and they stay in your own store. Up for it this week?

landing page

Your agent sessions, graded $49 founding quarter: weekly scorecards plus drills built from your own failure patterns Take one of 25 first-cohort seats, starts October 15

deposit terms

$49 covers your first three months in the founding cohort of 25, starting October 15, and locks the founding rate when usage billing begins. Refund on request any time before your first scorecard is delivered. Your transcripts stay in a store you control; we only read the sessions you send.

Would you run this test?

One tap. The yes-share feeds the Demand pillar of this idea's score; nobody sees who answered.

Budgets are out-of-pocket estimates for a team of one to three, US market. Size the deposit to the deal, and check the terms before taking money in a regulated line.

Scorecard

Ranked against every idea in the catalog: trend, demand and 100x potential from the corpus, competition relative to the other ideas. A generated concept has no judges or swipes yet, so its pillars use the data signals only.

49

Idea Score, 0-100 (partial) · raw 30.5 x 1.61

Warm

competition: more crowded than 32% of ideas · headwind x0.84

+0.0

government priorities, secondary (0 matching grants)

Trend

41

Is the wave forming now? 2025-26 entrants vs 2023-24, rounds since 2025, the sector's live-batch direction, the 2026 trend analyst.

  • Entrants 2025-26 vs 2023-24 (similar companies)74
  • Rounds announced 2025+ in the sector0
  • Sector direction (live batch)50

Demand

23

Does anyone want it? YC's current RFS, companies already paid for something similar, the operator judge, founders' yes-rate in decks, readers who would run the test.

  • YC asks for it (current RFS: idea / sector)30
  • Someone already pays (similar companies, recent / all-time)17

100x potential

50

Can it return a fund? The venture judge (double weight), market-size and moat axes, neighbours still alive, the technologist judge.

  • no signal yet, taken as 50

Score = 100 x cbrt(Trend x Demand x 100x) x (1 - 0.5 x crowding) + government bonus (max 5), calibrated so the 95th-percentile idea scores 90 (order never changes). A geometric mean: a weak pillar cannot be papered over. Percentiles are among the 382 ideas in the catalog; the terms matched were mcp, server, scores, sessions, drills, weakest, plugs, whatever.

The concept in full

What
Postgame plugs into whatever assistant you already run via MCP, keeps your transcripts in a store you own, and grades each session on cost, verification, delegation and rework, then generates personalized drills from your actual failure patterns, the prompt you re-ran four times, the output you accepted unverified. A daily AI user installs the server, replays last week's sessions, and sees their first scorecard within the hour; pricing is by usage.
Grounded in (2025-2026 signals)
MCP donated December 9, 2025 with 97 million monthly SDK downloads and 10,000 active servers; Claude Max launched April 9, 2025 at $100 and $200 a month for people who 'collaborate with Claude extensively'; Woz (YC W25) cuts Claude Code token cost 50% via plugin, proving session-level waste is real and measurable; MIT NANDA (August 2025): users call the same models reliable in their own hands and unreliable in enterprise systems.
What it rides
Rides 'MCP donated to the Agentic AI Foundation' (December 9, 2025): with 97 million monthly SDK downloads and first-class support in ChatGPT, Claude, Cursor, Gemini, Copilot and VS Code, one server can observe a user's practice across every assistant they run, which is what makes coaching from your own data possible.
Why now
The $100 and $200 Max tiers that opened April 9, 2025 created users spending $1,200 to $2,400 a year on capacity with zero feedback on how well they use it, and MCP's December 9, 2025 standardization means one integration now reaches every major assistant instead of six.
Wedge: first customer and entry point
First customers are Max and ChatGPT Pro subscribers who already track their token spend; the entry point is a free week of scored replays showing dollars wasted on rework, then usage pricing that stays a fraction of the waste it removes, buildable by one person as a single MCP server.
Closest real companies, as the generator saw them
Woz reduces token cost automatically inside Claude Code but teaches nothing and covers one tool; Confident AI observes LLM apps for engineering teams, not individuals. jo is a personal agent that does work rather than coaching how you work. None score the human across assistants.
Main risk
Assistants start shipping native usage analytics and self-critique, collapsing the space between the model and the coach.

Similar startups in the directory

Companies whose pitch matches most of the concept's terms (mcp, server, scores, sessions, drills, weakest, plugs, whatever).

  • AgentCatspeedrun SR007 · 2026 · Agent infrastructurealive

    See exactly how agents experience your product and where they get stuck.

  • Cosmicyc W19 · 2019 · Developer toolsalive

    Cosmic is an AI-powered headless CMS

Public money in this direction

US federal grants and open opportunities matched to the concept's terms.

No matching grants or programs tracked; 19 startup-relevant grants exist in Horizontal AI assistants overall.

Other concepts in this collection

Fictional concept generated 2026-08-26 by claude-fable-5 from the collection's brief and MarkosWeb data. Treat it as a research prompt, not a plan.