All topics

inference-engine

Projects tagged with inference-engine on GitHub.

16 projects

ds4 logo

ds4

C
95

DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm

21.9k+2260Jan 21, 1970
tokenspeed logo

tokenspeed

Python
90

TokenSpeed is a speed-of-light LLM inference engine.

2.0k+560Jan 21, 1970
open-webui logo

open-webui

Python
100

User-friendly AI Interface (Supports Ollama, OpenAI API, ...)

150.6k+9040Jan 21, 1970
MTPLX logo

MTPLX

Python
90

3x faster speeds on MLX | Qwen 3.8 27B | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.

1.9k+4820Jan 21, 1970
shard logo

shard

Python
72

Pipeline-parallel LLM inference across GPUs on separate machines.

447+40Jan 21, 1970
inference-school logo

inference-school

Swift
77

A hands-on Swift and Metal course for building LLM inference from first principles on Apple silicon, with 48 guided lessons, runnable exercises, a native macOS Studio, and a complete companion book.

185+10Jan 21, 1970
apex-inference-chip logo

apex-inference-chip

Python
68

An inference chip design that runs a real LLM (Qwen2.5-0.5B) on FPGA — one transformer decoder layer in RTL, every silicon value bit-exact against a golden model. 0.56 tok/s measured, a 140× climb, full evidence trail.

661+490Jan 21, 1970
llama.cpp logo

llama.cpp

C++
95

LLM inference in C/C++

126.4k+1.1k0Jan 21, 1970
bindwidth logo

bindwidth

JavaScript
72

Evidence-aware on-prem LLM inference sizing and TCO calculator

123+30Jan 21, 1970
cliare logo

cliare

Rust
72

CLI agent-readiness measurement, command-shape inference, and CI scorecards

4960Jan 21, 1970
transformers logo

transformers

Python
95

🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.

164.7k+2690Jan 21, 1970
Audar-ASR-V1 logo

Audar-ASR-V1

Python
68

Arabic-first generative speech recognition — Audar-ASR-V1 (Flash + Turbo). #1 on the Open Universal Arabic ASR Leaderboard. Model cards, benchmarks & inference.

4980Jan 21, 1970
ollama logo

ollama

Go
95

Get up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.

179.9k+5300Jan 21, 1970
local-llm logo

local-llm

Shell
70

Everything I know about running LLMs locally

1.8k+200Jan 21, 1970
kimodo.cpp logo

kimodo.cpp

C++
72

Animate skeletons with natural language; NVIDIA's Kimodo ported to C++/GGML

5310Jan 21, 1970
hayamimi logo

hayamimi

Python
72

早耳 - Real-time multilingual speech-to-text on CPU only. Live subtitles, browser dashboard, speaker labels, translation. No GPU, no cloud.

3070Jan 21, 1970