Skip to content

AI model

nomic-embed-text

Text embedding model used to compute cosine similarities between the twenty Spanish question pairs.

Current stories

buildOne report1 publisher

Ollama's timings block rates a one-token prefill at 3,875 tokens per second

Ollama's OpenAI-compatible endpoint reported 3,875 prompt tokens per second where the real rate was 125, a dev.to author found. Any benchmark that repeats a prompt and reads that field overstates prefill speed by the share served from cache.

Publishers:dev.to

Reality

Evidence55
Adoption
Insufficient
Hype gap+5
Incentives35
Confidence55