Solutions & capabilities

Systems built and operated for real work

Each entry is something I have built, run, and can defend with evidence: live URLs, write-ups, or open source, with a real last-updated date. The complete inventory is date-ordered within each grouping.

Earlier products and experiments live separately in Earlier products and experiments, with what each one taught.

AI & Agent Systems

Systems that keep agents running, measured, and honest about cost.

Leaderboard diagram showing browser automation tools ranked by median task time, with a second ranking for real-work capability.
AI & Agent Systems

Web Automation Microbench

Head-to-head benchmarks of 33 browser automation tools on one identical task, with a second real-work ranking that reorders them.

PythonOpenRouterChrome DevTools Protocol
Coverage matrix showing models along the rows and agent benchmarks along the columns, with cells coloured by provenance class and left grey where no result has been ingested.
AI & Agent Systems

Agent Benchmark Matrix

A provenance-first coverage matrix of agent, computer-use and tool-use benchmarks, with harness separation and honest gaps.

Next.jsTypeScriptVitest
LLM Pareto Frontier chart plotting model quality against cost, with efficient frontier models connected across the upper-left edge.
AI & Agent Systems

LLM Pareto Frontier

A live quality-versus-cost map of major language models, with the efficient frontier calculated in the browser.

Next.jsTypeScriptECharts
Control room architecture diagram showing agent telemetry flowing into a local Grafana operations room with alerts and runbooks.
AI & Agent Systems

Agent Operations Control Plane

Launchd-managed workers, an ops board, alerts, and runbooks for a personal agent fleet.

launchdSemaphoreAlertmanager
Diagram of one agent routing work between a cheap workhorse model and a frontier model.
AI & Agent Systems

Agent Routing and Lifecycle System

Role-aware routing from the cheap workhorse to frontier help, with A2A as a discovery capability.

CodexOpenCodeA2A
Session token shape diagram showing token usage across a coding agent session.
AI & Agent Systems

Coding Agent Observatory

Token, cost, latency and session telemetry across every coding agent I run.

PythonOpenTelemetryPrometheus
Routing benchmark architecture diagram showing identical tasks measured across provider policies with telemetry.
AI & Agent Systems

Model Routing Performance Lab

Benchmarked provider routing against measured latency, cost and quality.

OpenRouterbenchmarkingGLM-5.3-Flash
Local LLM Lab interface showing the installed-model inventory and measured comparisons.
AI & Agent Systems

Local LLM Lab

A measured field guide to every local model installed on Apple Silicon.

PythonReactVite
Model Intelligence Maintainer interface comparing model presets and recommended fits.
AI & Agent Systems

Model Intelligence Maintainer

A workbook and guide comparing model quality, price and provider coverage.

PythonOpenRouterArtificial Analysis
Agent Orchestra demonstrator interface showing orchestrated agent interaction.
AI & Agent Systems

Agent Orchestra

A demonstrator for orchestrated multi-agent interaction patterns.

agentsorchestrationNext.js

Martech & Measurement

Governance, QA, and reconciliation that make marketing data trustworthy.

AI Discovery Intelligence observation plane showing a platform comparison with evidence-backed fields, an explicit unknown gap and a sourced claim drill-down.
Martech & Measurement

AI Discovery Intelligence

A source-backed view of how consumer AI surfaces find, retrieve, cite and recommend brands — and which marketing actions the evidence actually supports.

Next.jsFastAPIPostgreSQL
A qualified planner answer showing the conditional outcome, its optimisation-signal control mode, prerequisite and evidence date, above a master table of capabilities with control, availability and evidence-basis columns.
Martech & Measurement

Ad Platform Capability Explorer

Qualified planner answers for advertising-platform capabilities: supported / conditional / unknown, with conditions, evidence basis and a verification date, over a searchable master table of the published corpus.

Next.jsTanStack TableTanStack Virtual
Agentic analytics loop diagram showing data flowing through collection, warehousing, QA, and reporting stages.
Martech & Measurement

Global Measurement Governance System

Taxonomy, QA, and reconciliation for measurement across markets and vendors.

GA4GTMBigQuery
Source-mix calibration diagram comparing observed channel data with platform-reported and reconciled measurement.
Martech & Measurement

Media QA and Attribution Reconciliation Toolkit

Tag QA, pixel checks, and attribution reconciliation that survive consent and real browsers.

browser QAtag QAattribution
Open GTM Index interface showing category leaders and ranked open-source tools.
Martech & Measurement

Open GTM Index

A transparent guide to open-source go-to-market software.

TanStack StartReact 19TypeScript

Products & Operational Tools

Working products that support real operations, not demos.

Knowledge systems loop diagram showing source material feeding a maintainable knowledge layer and review gates.
Products & Operational Tools

AI-Assisted Product Definition System

Turns messy commercial context into PRDs, plans and scoped decisions.

PRD workflowplanning flowsknowledge layer
Creative Observatory workbench interface showing brand controls and the evidence review workflow.
Products & Operational Tools

Creative Observatory

A source-aware workbench for reviewing public ad-library evidence.

Next.js 15PrismaRecharts
Hackathon Voting App interface showing the public scoreboard and consent-aware controls.
Products & Operational Tools

Hackathon Voting App

A production-ready single-screen judging app built for a live hackathon room.

Next.js 14ClerkPrisma