Skip to content

Topic

Synthetic Training Data

Generation and use of model-produced data for pre-training and fine-tuning, and the debate over its limits.

Current stories

product3 publishers

Salesforce's Koa model reaches pilots with two accounts of what it was trained on

Salesforce and Nvidia post-trained Koa on simulated CRM workflows and claim three times fewer errors than leading general-purpose models. The published record describes the training data two different ways.

Perspective Coverage

3 publishers
Builder
Builder 35%
Operator
Operator 38%
Investor
Investor 27%

Reality

Evidence45
Adoption20
Hype gap+35
Incentives75
Confidence55

Earlier coverage

  1. Robot suppliers tell an Ulsan forum that plants allow testing and ban filming

    Invest · September 12, 2026 · 1 publisher

  2. Generated tree images help a species classifier only where real photographs are thin

    Science · September 11, 2026 · 1 publisher

  3. Google's ToolGrad back-writes the user query from a chain it has already executed

    Build · September 10, 2026 · 1 publisher

  4. Escalating retries recover the 7,930 refusal prompts a single steering pass drops

    Build · September 8, 2026 · 1 publisher

  5. Darkening the skin around a benign mole flipped GPT-4's answer to melanoma

    Product · September 6, 2026 · 1 publisher

  6. A 1.1-second fixed overhead pushes caption-driven TTS out of the conversational path

    Build · August 30, 2026 · 1 publisher

  7. The judge went synthetic first, which tells you which part of your pipeline is next

    Build · August 22, 2026 · 1 publisher

  8. Picovoice's free tier is gone. Re-cost the always-on layer at 53 ms per second.

    Build · August 21, 2026 · 1 publisher

  9. Sutton calls synthetic data 'a big mistake': every simulator is a lossy copy of a bigger world

    Product · August 19, 2026 · 1 publisher