Skip to content

Topic

Multimodal inference serving

Serving models that accept images, video or audio alongside text, where media must be turned into embeddings by an encoder before language-model prefill can begin.

Current clusters