evals
Projects tagged with evals on GitHub.
9 projects
skillspec
RustSkillSpec makes agent skills followable, testable, and provable with Doctor risk reports, guided imports, structured contracts, and alignment proof.
waku-agent
PythonWaku Waku! Waku agent is your personal AI agent, on your own laptop, in code you can read in an afternoon — harness + loop + memory + eval
PerceptionBench
PythonPerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
DeepSpec
PythonDeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms
llm-space
TypeScriptA desktop app to prototype agent ideas, inspect every harness step, replay failures, and evaluate performance, all in one place. Local-first, cloud-ready for managed agents.
agent-apprenticeship
PythonThe living ecosystem where AI agents complete tasks through workflow loops, improve through iterative execution, are evaluated by mentor agents or humans in the loop, and turn completed work into reusable work experience and data to improve future agents.
agent-apprenticeship
PythonThe living ecosystem where AI agents complete tasks through workflow loops, improve through iterative execution, are evaluated by mentor agents or humans in the loop, and turn completed work into reusable work experience and data to improve future agents.
forsy-trace-skill
PythonOpen skill for capturing AI agent work as structured traces.
better-harness
JavaScriptBetter Harness turns project and session evidence into loop-level insights, prioritized improvements, and verifiable next steps—inside the coding agent you already use.