TESIGN
SAVED

EN · KO

  1. TESIGN / RADAR
  2. 2026.09.22
  3. AgentMeasure

AgentMeasure

Audit your Codex and Claude Code logs

WHO USES ITDevelopers who use Codex or Claude Code daily, and anyone building or maintaining agent tooling.

Install to use

Reads your Codex and Claude Code session logs and surfaces repeated failures and retries, with the evidence attached.

INSTALL
START IN 5 MINUTES ↓
The AgentMeasure landing page: "The open yardstick for agent metrics" on the left, and a terminal panel on the right showing checks marked PASS, FAIL and UNPROVABLE.

TESIGN TAKE

Marking what it cannot prove as UNPROVABLE instead of quietly writing zero is what makes the numbers worth trusting.

Activity and discovery evidence
★ 216▲ 2 HN

216 stars · checked on GitHub GitHub check 68 h late · no 7-day star increase observed (GH Archive) (as of ) · 2 Show HN points observed

NEW 0h OLD WHEN SEEN

AT A GLANCE

LICENSE
MIT
USAGE
Use, change and redistribute, commercially too. Keep the notice.
LANGUAGE
Python
PLATFORM
cli
ACTIVITY
last commit 2 days ago () · latest release [unconfirmed]
COMMUNITY
contributors [unconfirmed] · open issues [unconfirmed] · made by: [unconfirmed] Reference time unavailable
SOURCES
Show HN
OPEN SOURCE
YES
FIRST SEEN
CATEGORY
DEV TOOLS · AI

A summary, not legal advice.

WHY IT MATTERS

An agent that edits the same file over and over, or hits one failed tool call after another, is easy to miss when it is buried in a log. AgentMeasure reads the session logs you already have and finds duplicate records, retry chains and consecutive tool failures — and when the evidence is not there it says UNPROVABLE rather than inventing a zero.

BUILD FROM THIS

  • Written in Python and published on PyPI, so pipx runs it with no repository checkout. There is a demo mode with no personal logs needed, a check command over the last seven days, and a flag to force the Claude Code adapter. All analysis runs locally with no network calls, and the result is a terminal summary plus a local HTML report.

WHO IT'S FOR

Developers who use Codex or Claude Code daily, and anyone building or maintaining agent tooling.

START IN 5 MINUTES

# Run pipx run agentmeasure demo first to see it on a synthetic example.

CAVEATS

  • It is built around the log formats Codex and Claude Code write, so sessions from other tools are not covered, and the repository marks its billing/settlement feature (settle) as still a concept.

RECEIPT

FIRST SEEN
AT SOURCE
KEPT
CREATED → FIRST SEEN
0h
WHEN FIRST SEEN
216 ★
SOURCES
Show HN

The same facts in machine-readable form — View as Markdown · JSON

SIMILAR TOOLS

Up to five from DEV TOOLS by ★ total: edited entries first, then repository cards; this entry is highlighted. Drawn from the same stored snapshot as the rankings — a comparison, not a recommendation.

IMAGENAME★ TOTAL▲ 7dLICENSEPLATFORMLAST PUSH
ECC README banner: headline 'The operating system for AI agent harnesses', harness chips, and skill, agent and command lists.
eccSkills, memory and security for agents★ 265k▲ +413MITpermissivewindows · macos · linux · cli
DeepSeek Harness landing: blue background, the headline 'Everything is a plugin' and a quick-start box with the npx command.
DeepSeek HarnessPlugin agent harness for developers★ 232k▲ +487MITpermissiveweb · cli
ponytail.dev landing page: a line-drawn developer with a ponytail on black, the title 'ponytail' and an install button.
ponytailMake your agent write less code★ 144k▲ +349MITpermissive
Spec Kit docs landing page — the headline 'Build with a spec, fix a bug, or assess an idea' and the uv tool install specify-cli command
Spec KitSpec-first process for coding agents★ 138k▲ +89MITpermissivecli · windows · macos · linux
The AgentMeasure landing page: "The open yardstick for agent metrics" on the left, and a terminal panel on the right showing checks marked PASS, FAIL and UNPROVABLE.
AgentMeasureAudit your Codex and Claude Code logsTHIS ENTRY★ 216MITpermissivecli

TIMELINE

  1. Repository created
  2. Show HN post ↗
  3. FIRST SEEN BY TESIGN
  4. Published on TESIGN

SAME DAY

  1. 01

    ZCode

    Run multiple coding agents from one app

    A coding workspace that runs several AI agents at once and drives a goal to completion.

    Use in the browser
    ★ 5.1k▲ +64 / 7d
  2. 02

    Try Omarchy for Windows

    Try the Omarchy Linux desktop on Windows

    An app that boots the real Omarchy Linux desktop inside a window on Windows — no VM software to install.

    Install to use
    ★ 440▲ 5 HN
  3. 03

    hibi

    A markdown notes app with a graph view

    A desktop Markdown editor that also draws a graph of how your notes connect.

    Install to use
    ★ 14▲ 5 HN
  4. 04

    tfm

    Drag and drop files inside your terminal

    A terminal file manager built for the mouse — drag, drop, dual panes, thumbnails.

    Install to use
    ★ 136▲ +9 / 7d
  5. 06

    package-doctor

    Find the Python dependency to fix first

    Scans Python dependencies and tells you which are being exploited now — and which have no one left to fix them.

    Install to use
    ★ 8▲ 4 HN
  6. 07

    pokertools-arena

    Watch AI models play poker at one table

    Seats several OpenAI-compatible models at the same poker table and lets you watch every decision as a spectator.

    Use in the browser
    ★ 1▲ 2 HN
  7. 08

    Roastbook

    Log espresso shots on your own server

    A self-hosted coffee journal that reads a photographed bag’s label and fills in the details for you.

    Self-host
    ★ 2▲ 1 HN
  8. 09

    Comic Chat for AI

    Turn a ChatGPT chat into a comic strip

    A revival of Microsoft’s old Comic Chat that automatically draws a ChatGPT conversation as a comic strip.

    Install to use
    ★ 1▲ 3 HN
  9. 10

    2049 — Mines of Merge

    2048 fused with Minesweeper in a dungeon

    You are the tile: slide by 2048 rules while dodging Minesweeper-style hidden mines down eight depths.

    Use in the browser
    ★ 0▲ 1 HN