All tools › Models & AI Labs › Cerebras Inference Cerebras Inference Models & AI Labs Fastest LLM inference. Llama 3.3 70B at 1000+ tok/s. Free tier. 🌐 Visit website ↗ 🔗 Similar tools Groq Cloud Ultra-fast LPU inference. Mixtral, Llama, Gemma. Free API tier. OpenAI Jalapeño OpenAI Jalapeño is OpenAI's custom inference chip, with first public results from August 2… Together AI 200+ open models. Fast inference API. Free tier. llama.cpp ★ 127.1k llama.cpp is a C/C++ inference engine for large language models that runs on CPU, GPU, and… DeepSeek V3/R1 ★ 104.4k DeepSeek V3/R1 is a 671B-parameter mixture-of-experts model offering GPT-4-level reasoning… GPT4All ★ 77.4k GPT4All is an open-source local chat application that runs large language models directly …