Create Alert
Email me similar jobs

Runtime Engineer - Open Models, GPU Backends & Performance

Ollama is the leading open-model runtime that runs locally on developers' machines, supporting macOS, Linux and Windows across diverse hardware. You’ll work on the heart of Ollama — the runtime that loads models, manages memory, and drives GPU acceleration for NVIDIA, AMD, Intel, Qualcomm, and Apple Silicon.

You’ll write in Go and C/C++ and touch model formats, inference engines, and the orchestration that makes open models feel instant on a laptop or server.

#J-18808-Ljbffr
Similar jobs

Runtime Engineer - Open Models, GPU Backends & Performance

Apply Now
Back to search page