inference-engine
Projects tagged with inference-engine on GitHub.
16 projects
ds4
CDeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
tokenspeed
PythonTokenSpeed is a speed-of-light LLM inference engine.
open-webui
PythonUser-friendly AI Interface (Supports Ollama, OpenAI API, ...)
MTPLX
Python3x faster speeds on MLX | Qwen 3.8 27B | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.
shard
PythonPipeline-parallel LLM inference across GPUs on separate machines.
inference-school
SwiftA hands-on Swift and Metal course for building LLM inference from first principles on Apple silicon, with 48 guided lessons, runnable exercises, a native macOS Studio, and a complete companion book.
apex-inference-chip
PythonAn inference chip design that runs a real LLM (Qwen2.5-0.5B) on FPGA — one transformer decoder layer in RTL, every silicon value bit-exact against a golden model. 0.56 tok/s measured, a 140× climb, full evidence trail.
llama.cpp
C++LLM inference in C/C++
bindwidth
JavaScriptEvidence-aware on-prem LLM inference sizing and TCO calculator
cliare
RustCLI agent-readiness measurement, command-shape inference, and CI scorecards
transformers
Python🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Audar-ASR-V1
PythonArabic-first generative speech recognition — Audar-ASR-V1 (Flash + Turbo). #1 on the Open Universal Arabic ASR Leaderboard. Model cards, benchmarks & inference.
ollama
GoGet up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
local-llm
ShellEverything I know about running LLMs locally
kimodo.cpp
C++Animate skeletons with natural language; NVIDIA's Kimodo ported to C++/GGML
hayamimi
Python早耳 - Real-time multilingual speech-to-text on CPU only. Live subtitles, browser dashboard, speaker labels, translation. No GPU, no cloud.