Abundant

ActiveAI-native

Agent simulation and RL for researchers

What it does

Hello! 👋 We are a team of former ML engineers, founders, roboticists and ops leads who obsess about data and its impact on safe, reliable AI. We specialize in creating environments and datasets for RL by leveraging our experience in simulation and model training. By the numbers: • Powering 3 of the top 6 global AI labs and multiple Fortune 500 enterprises • Billions of training tokens generated each month, 2x month over month • Exclusive, global network of over 500 domain experts We believe humans are inherently creative, and thrive by pushing the frontier. We are working towards an abundant future--one where everyone has access to infinite intelligence, services and goods. Based in San Francisco, CA. We enjoy good food and good company. -- more info below -- Abundant is building the NVIDIA of training data. AI models rely on two fundamental ingredients: compute and data. NVIDIA, the leader in compute, has a peak market cap of $5T and generated $130B in revenue last year as the need for scaling compute has exploded. We believe the need to scale data is just beginning, as we move beyond SFT and human supervision to RL and Learning from Experience. Our founding team consists of second-time founders, ML engineers and data leads from Waymo, Google, Meta and AWS. Our team has previously collaborated with DeepMind to classify hate speech in YouTube videos, trained SOTA models for self-driving, and scaled data pipelines with thousands of human annotators. Our pioneering work in human computation, synthetic data, imitation learning and RL give us a solid advantage in delivering results to our customers. Why now? Training data is more important and more scarce than ever before. Scaling laws dictate that linear improvement in model performance demands an exponential increase in training data. But there is only one World Wide Web and most of it has already been trained on. The next advances will require new, diverse, and high-quality dat

Its site says now · captured 2026-08-23

abundant

Environments and datasets for RL. Powering the world

Next to it in Data for AI · RL simulation and synthetic data generation

  • NeosyncY Combinator S22

    Neosync is an open-source anonymization and synthetic data platform.

  • Synthetic SocietyY Combinator S25

    Synthetic Users to Simulate Real Users

  • BesampleTechstars TS 2024

    Besample | Research beyond the West

  • DatacurveY Combinator W24

    Frontier coding data for training and evaluating LLMs

  • PerfectBit, Inc.Y Combinator X26

    Correct by construction AI training data

  • Olam LabsY Combinator S26

    Building multi-agent simulations for model evals and training.

  • QokedasY Combinator F26

    Data for AI Science

  • Sepal AIY Combinator S24

    Data Development for Advanced AI