Runworthy
The agent operations scanner: pip install runworthy, point it at a repo, read the graded report — what's confirmed, what to verify, what only you can answer.
Phase 1 shipped — scans now come back graded. Next: the web report.
Runworthy reads your repo, asks the questions code can't answer, and gives a plain verdict: GO or NO-GO, and what to fix first. The framework behind it is free for anyone.
The agent operations scanner: pip install runworthy, point it at a repo, read the graded report — what's confirmed, what to verify, what only you can answer.
Phase 1 shipped — scans now come back graded. Next: the web report.
A progressive deep-research engine that plugs into any MCP client — persistent knowledge graph, source trust scoring, and it tells you when to stop. 78 tests, dogfood-validated.
The rubric Runworthy grades against: 29 operational controls across six domains, ten of them non-negotiable. Free under CC BY 4.0, and written to sit alongside OWASP's agentic Top 10.
I build agent systems on Firebase and Google Cloud at Obsidicore, and safety tooling for anyone else running agents. Before that I led engineering at an EdTech company, shipping RAG pipelines for K-12 school districts on Vertex AI.
The Air Force is where I learned the operational side: security operations, contingency planning, checklists written for the day something breaks. Lava was my callsign. It stuck to the domain, and the rest shows up in how I build.
I think in systems, not silos. Father of three.
Two companies' worth of things I'd do differently, written down while they still sting.
Architected multi-tenant infrastructure for dozens of school districts when we had three. That engineering time should have gone to features that drove adoption.
Validate demand before scaling infrastructure.
20+ CLI commands and four content pipelines before the strategy was validated. A pivot wrote most of it off.
Automate after the workflow is proven, not before.
Range is my strength for novel problems, but it can underestimate depth work. Now I pair with specialists instead of fumbling through alone.
Know when to bring in experts.
900+ assets/month across four automated pipelines — Gemini analysis, summarization, knowledge-graph construction on Cloud Functions.
Multi-brand publishing pipeline: X API distribution, Cloud Scheduler timing, 20+ operational CLI commands.
Curriculum processing on Vertex AI embeddings + vector search; multi-tenant Firestore serving school districts.
MCP research engine — knowledge graph, trust scoring, self-regulation. MIT.
Most of what I make ends up on GitHub, mistakes included.
I take a few Firebase, GCP, and agent-safety engagements a year.