Skip to content

Topic

GPU acceleration

Moving computation from the CPU to a graphics processor. For language model inference it turns on whether weights and cache fit in video memory.

Current clusters