Wafer

ActiveAI-native

AI that makes AI fast

What it does

Wafer builds AI agents that work as autonomous performance engineers, optimizing GPU kernels for AI inference. Our product is serverless and dedicated inference for the world’s fastest open source LLMs, achieved by Wafer's autonomous performance engineers.

Its site says now · captured 2026-08-23

Wafer | LLMs for enterprise

The fastest inference for open models.

Next to it in AI infra and compute · GPU kernel optimization for inference

  • deepsiliconY Combinator S24

    Software and hardware to run neural networks faster and cheaper

  • RiftenY Combinator S26

    Earned intelligence for every company

  • LuminalY Combinator S25

    Making AI run fast on any hardware.

  • BosonicSOSV SOSV HAX Seed 2025

    Building the Engine of the Imagination Age where brilliant people and limitless compute condense to transform the world.

  • TracerY Combinator S26

    Combining open-source AI models for better answers at lower cost

  • Understudy LabsY Combinator S26

    Self-optimizing neocloud that cuts LLM bills by 80%

  • RightNowY Combinator F26

    Enabling Model-Hardware Co-Design at Scale

  • DownlinkY Combinator W24

    Make your LLMs 3x faster.