Projects

Built, not pretended.

Most of the work is held in confidence. What can be shown sits here, framed for the question that matters to a business: what it changes, and where the judgment stays human.

9 shown· 1 live trace· 5 case study· 2 public artifact· 1 private capability
Public · Trace Proof of Work · live 9.95B tokens 169 days operating since 27 Jan 2026 · peak 596.8M in one day · adversarial across Claude, Codex, Hermes, Grok build CLI and OpenCode, with OpenClaw as the retired predecessor · read locally, nothing uploaded. The number is a trace, not a trophy: the page shows where the spend creates value, waste, or risk. Claude Code prunes old logs, so its total is held at the highest figure ccusage ever measured, its June 24 record plus everything burned since. Grok's number is a floor from its surviving log. OpenClaw operated from Feb 3 to May 30, 2026, and its 44.3M is estimated from recovered logs plus its confirmed daily operating history. Open the live dashboard →
Workshops & Curricula
Curriculum · 20 modules
Public · Case study
Enablement // Production readiness // Evaluation

AI Evaluation Workshop

For teams who need AI that ships, not demos. Twenty modules on evals, datasets, guardrails, red teaming, and the quality gates that keep a model honest in production.

ProofA complete, teachable evaluation method, not a single eval script: each module is a decision a team has to settle before it can trust a model in production.
StackEvals · datasets · guardrails · red teaming · quality gates
Why it mattersMove from an impressive demo to AI you can actually trust in production.
Evaluation
Workshop · Leaders
Public · Artifact
Enablement // Leadership // Systems thinking

AI Systems Thinking Workshop

Teaches leaders without a technical background to read their own operation like a systems engineer, so they can see where AI belongs before they buy it.

ProofA reusable teaching method that gets non technical leaders to map their own operation and locate where AI fits, before any tool is bought.
StackSystems mapping · process decomposition · adoption sequencing
Why it mattersDecide where AI belongs before spending on tools that do not fit the operation.
View on GitHub →
Systems Thinking
Voice & Conversational AI
Production Playbook · 39 modules
Public · Case study
Voice AI // EU compliance // Production playbook

Voice Agent Deployment Kit

A production playbook for shipping voice agents that survive contact with EU regulation. Covers SLAs, GDPR and the AI Act, vertical scripts, and the adoption maps most teams discover too late.

ProofMaps what it actually takes to ship a voice agent under EU regulation across 39 modules, so the compliance and adoption work is planned before launch instead of discovered after it.
StackSLAs · GDPR · EU AI Act · vertical scripts · adoption maps
Why it mattersShip a voice agent in the EU without a compliance surprise after launch.
Voice Agents
Open Research
Co-authored preprint · Zenodo
Public · Case study
Co-authored research // AI memory // Provenance

MaluDB: Long Term AI Memory

Organizations store records, documents, tickets, and messages, but lose the evidence, timing, permissions, and reasoning that make that knowledge safe to reuse. MaluDB proposes a memory-native database for that missing layer.

ProofA co-authored open preprint on Zenodo that specifies a memory-native DBMS architecture for institutional memory shared by humans, applications, and AI agents.
StackMemory objects · provenance · temporal validity · permissions · audit trail
Why it mattersThe difference between AI that compounds what your teams know and AI that repeats expensive mistakes.
MaluDB
Tools & Systems
Open Skill · Prompt System
Public · Case study
Open skill // Interview prep // Discovery prep

Call Readiness Coach

The first Systems Detective open skill. A copy-paste prompt that pressure tests your preparation for the conversations that decide things: job interviews, discovery calls, and client onboarding. One evidence-grounded engine, three modes.

ProofAn evidence first preparation drill that asks one question at a time, names the weak spot, and refuses the comfortable hero story, with the full prompt public and installable as a skill.
StackCopy paste prompt · Claude skill · Codex skill · three modes
Why it mattersWalk into a high-stakes conversation with an answer that survives the second question, not a story that falls apart.
Call Readiness
Healthcare privacy · Open source
Public · Case study
Healthcare // Privacy // Data governance

Healthcare Data Privacy Lab

An offline anonymisation prototype and open-source contribution track for safer healthcare data workflows. I first built a local health-note anonymiser because sensitive health data should not move into AI tools, SaaS products, or shared workflows by default. I later found OpenMed and started contributing to the larger ecosystem.

ProofBuilt a public offline anonymisation prototype around local data minimisation, then joined OpenMed's broader healthcare privacy and reliability work.
StackHIPAA Safe Harbor · GDPR · EU AI Act · local-first systems · de-identification · OpenMed · Python · clinical text · data governance
Why it mattersHealthcare systems should not start by uploading raw patient text. They should start by removing what the external system does not need.
Privacy Lab
Available on request
Private · Capability
Due diligence // Verification // Risk

OSINT Due Diligence

Tooling to verify who you are actually dealing with before a deal, a hire, or a partnership. Evidence first, assumptions last.

ProofA working capability run before a deal, a hire, or a partnership: it establishes who is actually on the other side from open evidence, and it is delivered privately rather than as a public download.
StackOSINT · identity verification · evidence trail
Why it mattersKnow who is on the other side of a deal, a hire, or a partnership before you commit.
Start a conversation →
OSINT
Discipline Tool
Public · Artifact
Discipline // Focus // Execution

Anti Rabbit Hole

A discipline tool that catches research spirals and redirects to execution, because the most expensive procrastination is the kind that looks like work.

ProofA working check that catches a research spiral and redirects it to execution, so the most expensive procrastination, the kind that looks like work, gets interrupted.
StackDiscipline prompt · spiral detection · redirect to execution
Why it mattersTurn research time into shipped decisions instead of busywork.
View on GitHub →
Anti Rabbit Hole

Most client engagements stay private by default. If a specific capability is relevant to your situation, it is shown and discussed directly. Start a conversation →