insideLLMs

Stop shipping LLM regressions. Deterministic behavioural testing that catches breaking changes before they reach production.

graph LR
    Dataset[Dataset] --> Runner[Runner]
    Model[Models] --> Runner
    Probe[Probes] --> Runner
    Runner --> Records[records.jsonl]
    Records --> Summary[summary.json]
    Records --> Report[report.html]
    Records --> Diff[diff command output]

The Problem

You update your LLM. Prompt #47 now gives dangerous medical advice. Prompt #103 starts hallucinating. Your users notice before you do.

Traditional eval frameworks can’t help. They tell you the model scored 87% on MMLU. They don’t tell you what changed.

The Solution

insideLLMs treats model behaviour like code: testable, diffable, gateable.

insidellms diff ./baseline ./candidate --fail-on-changes

If a configured behavioural gate fires, the deploy blocks.


Start Here

Goal Path Time
See it work Quick Install → First Run 5 min
Compare models First Harness 15 min
Block regressions CI Integration 30 min
Add provenance checks Verifiable Evaluation 15 min
Understand the approach Philosophy 10 min

The portable offline path is:

insidellms init harness.yaml --template harness
insidellms harness harness.yaml --dry-run
insidellms harness harness.yaml --run-dir ./runs/baseline

The initializer creates both the config and its sample dataset, so this works after a package install as well as from a repository checkout.


Why Teams Choose insideLLMs

Catch Regressions Before Production

Know exactly which prompts changed behaviour. No more debugging aggregate metrics.

CI-Native Design

Built for reviewing model behaviour as versioned artifacts. Stable artifact contracts and automated gates make behavioural changes visible; hosted-model responses can still be stochastic.

Response-Level Visibility

records.jsonl preserves every input/output pair. See what changed, not just that something changed.

Provider-Agnostic

OpenAI, Anthropic, Cohere, Google, local models (Ollama, llama.cpp, vLLM). One interface.


How It Works

1. Define behavioural tests

probes:
  - type: logic      # Reasoning consistency
  - type: bias       # Fairness across demographics
  - type: attack     # Jailbreak resistance

2. Run across models

insidellms harness config.yaml --run-dir ./baseline

3. Catch changes in CI

insidellms diff ./baseline ./candidate --fail-on-changes
# Exit code 2 for regressions, other changes, or one-sided records

Without a fail flag, diff is informational and exits 0 even when it reports differences. Improvements alone do not fail --fail-on-changes, and trace or trajectory findings require their dedicated gate flags.

Result: Breaking changes blocked. Users protected.


Documentation

Section Description
Philosophy Why insideLLMs exists and how it differs
Getting Started Install and run your first test
Tutorials Bias testing, CI integration, custom probes
Concepts Models, probes, runners, determinism
Advanced Features Pipeline, cost tracking, structured outputs
Reference Complete CLI and API documentation
Guides Caching, rate limiting, local models
FAQ Common questions and troubleshooting

What You Get That Others Don’t

Feature Eleuther HELM OpenAI Evals insideLLMs
CI diff-gating No No No Yes
Deterministic artefacts No No No Yes
Response-level granularity No Partial No Yes
Pipeline middleware No No No Yes
Cost tracking & budgets No No No Yes
Structured output parsing No No No Yes
Agent evaluation No No No Yes

Not just benchmarks. Production infrastructure.


Community