All topics

evals

Projects tagged with evals on GitHub.

9 projects

skillspec logo

skillspec

Rust
77

SkillSpec makes agent skills followable, testable, and provable with Doctor risk reports, guided imports, structured contracts, and alignment proof.

857+550Jan 21, 1970
waku-agent logo

waku-agent

Python
80

Waku Waku! Waku agent is your personal AI agent, on your own laptop, in code you can read in an afternoon — harness + loop + memory + eval

7360Jan 21, 1970
PerceptionBench logo

PerceptionBench

Python
78

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

1560Jan 21, 1970
DeepSpec logo

DeepSpec

Python
90

DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms

6.8k+3200Jan 21, 1970
llm-space logo

llm-space

TypeScript
90

A desktop app to prototype agent ideas, inspect every harness step, replay failures, and evaluate performance, all in one place. Local-first, cloud-ready for managed agents.

1.5k0Jan 21, 1970
agent-apprenticeship logo

agent-apprenticeship

Python
82

The living ecosystem where AI agents complete tasks through workflow loops, improve through iterative execution, are evaluated by mentor agents or humans in the loop, and turn completed work into reusable work experience and data to improve future agents.

1.3k0Jan 21, 1970
agent-apprenticeship logo

agent-apprenticeship

Python
82

The living ecosystem where AI agents complete tasks through workflow loops, improve through iterative execution, are evaluated by mentor agents or humans in the loop, and turn completed work into reusable work experience and data to improve future agents.

1.3k+90Jan 21, 1970
forsy-trace-skill logo

forsy-trace-skill

Python
67

Open skill for capturing AI agent work as structured traces.

970Jan 21, 1970
better-harness logo

better-harness

JavaScript
90

Better Harness turns project and session evidence into loop-level insights, prioritized improvements, and verifiable next steps—inside the coding agent you already use.

1.4k0Jan 21, 1970
evals Open Source Projects | MushyBook