Skip to content

project

llama-server

llama-server is the HTTP server component of llama.cpp, letting users run local LLM inference with configurable context, KV cache, and GPU settings.

Known aliases

  • llama.cpp server
  • llama serve

Relationships

No evidence-backed relationships are recorded.

Current stories

build1 publisherOne report

Speculative decoding in llama-server swaps real logprobs for 0.0 placeholders

llama-server b11430 reports logprob 0.0 for every speculatively decoded token, dragging one test's mean logprob from -0.48 to -0.0011. Nothing in the response or the server log flags the fill-ins, so evals and calibration built on those numbers go wrong quietly.

Publishers:dev.to

Reality

Evidence64
Adoption
Insufficient
Hype gap0
Incentives
Insufficient
Confidence58