Wafer
ActiveAI-nativeAI that makes AI fast
What it does
Wafer builds AI agents that work as autonomous performance engineers, optimizing GPU kernels for AI inference. Our product is serverless and dedicated inference for the world’s fastest open source LLMs, achieved by Wafer's autonomous performance engineers.
Its site says now · captured 2026-08-23
Wafer | LLMs for enterprise
The fastest inference for open models.
Next to it in AI infra and compute · GPU kernel optimization for inference
- deepsiliconY Combinator S24
Software and hardware to run neural networks faster and cheaper
- RiftenY Combinator S26
Earned intelligence for every company
- LuminalY Combinator S25
Making AI run fast on any hardware.
- BosonicSOSV SOSV HAX Seed 2025
Building the Engine of the Imagination Age where brilliant people and limitless compute condense to transform the world.
- TracerY Combinator S26
Combining open-source AI models for better answers at lower cost
- Understudy LabsY Combinator S26
Self-optimizing neocloud that cuts LLM bills by 80%
- RightNowY Combinator F26
Enabling Model-Hardware Co-Design at Scale
- DownlinkY Combinator W24
Make your LLMs 3x faster.