I'm building Shelfie, an iOS app written mostly by Claude Code agents, and the tests and CI that decide which of their work is safe to merge.

Every project below is a public repo with its tests and CI

notes on testing agent-written code

See the work

Work

9 shipped projects, every one with a public repo.

AI Infrastructure

LLM Gateway Kit

Semantic and vision response caching, tiered circuit breakers, and hard cost-budget enforcement for Swift LLM apps, built so a runaway prompt loop can’t blow through your API budget.

Problem: An iOS app calling LLM APIs directly has no natural circuit breaker: a retry storm, a redundant vision call, or a provider outage can burn through a monthly budget in hours, and usage dashboards only tell you after the money is gone.

Decisions:
  • Added semantic and vision response caching so near-duplicate prompts and images don't re-trigger a full paid call.
  • Built tiered circuit breakers that fall back tier by tier when a provider misbehaves, instead of a binary up/down switch.
  • Enforced cost budgets as a hard ceiling in the gateway itself, not as a downstream alert on a usage log, so the limit holds even if nobody is watching.

Outcome: A hard budget cap stops runaway spend before it happens. The patterns were extracted from the LLM gateway in an iOS app that is in TestFlight beta, and there is no published cost benchmark yet.

Swift · source

Verify Before Ship

A fact-checking gate for anything an LLM writes with citations: re-fetches every cited page and flags any claim it cannot find there, so a reviewer looks at those before anything is published.

Problem: Signal Scout's source-verification step proved the pattern worked for one skill, but every other project generating AI text with citations needed the same guarantee, and copy-pasting the check into each one meant fixing the same fabrication bugs repeatedly.

Decisions:
  • Generalized the containment-checking methodology out of signal-scout into a standalone tool, so any LLM-writing pipeline can adopt it as a dependency instead of reimplementing it.
  • Made it run before publication rather than as an audit afterwards: flagged claims go to a person to confirm or cut before anything ships.
  • Kept the check narrow and deterministic: re-fetch the source, confirm the claim is actually in it, rather than asking another model to grade the first model's honesty.

Outcome: Any project that generates cited claims can now pull in a reusable pre-publish citation check instead of bolting a one-off script onto a single skill.

Python · source

Scoped

An MCP server plus a PreToolUse hook that stops two Claude Code sessions from editing the same file at once. The hook blocks the edit itself, so it works even when the model never calls the tools.

Problem: Running a Claude Code session per issue or per worktree is normal now, and nothing stops two of them from editing the same file. An advisory lock API only helps if the model remembers to call it.

Decisions:
  • Enforced in a PreToolUse hook rather than relying on tool calls, and made the hook fail open so a bug in the coordinator never blocks real work.
  • Kept the lock in local SQLite and posted Linear comments only for visibility, since a network call with no compare-and-swap is the wrong place for the atomic part.
  • Tested the race with real OS processes sharing one database. Claims made one after another on a single connection passed; the multi-process tests exposed a check-then-insert race and a hook that ignored the insert result.

Outcome: In the process race tests, every round has exactly one winner and no errors. It has not been measured on a real multi-session fleet yet, and the README says so.

Node.js, SQLite, MCP · source

Developer Tools

Architecture Lint

A single bash script that checks module boundaries in TypeScript and Swift codebases, with a count-based baseline so you can turn it on before cleaning up existing violations.

Problem: Architecture rules are easy to agree on and hard to enforce retroactively: turning on a boundary linter for the first time on a real codebase means it immediately fails on years of pre-existing violations, so teams either skip enforcement entirely or burn a sprint on a big-bang cleanup before CI can go green.

Decisions:
  • Record a per-rule violation count at adoption time and fail CI only when a count rises. It counts rather than tracks individual violations, so fixing one and adding another in the same change still passes; I chose the simpler version and wrote that trade-off down.
  • Kept boundary rules in one list inside the script instead of per-file annotations, so adopting it doesn't touch application code.
  • Kept it to one bash script with no dependencies to install, and made it fail loudly if it's pointed at the wrong folder instead of silently reporting clean.

Outcome: You can turn on boundary checks the same day, on a codebase with existing violations, and CI stops the count from growing while you pay the debt down.

Shell, CLI · source

Litmus

Tests for prompt-based skills: deterministic checks wherever possible, and a model judge only where it has been checked against known good and bad examples.

Problem: Standard unit tests assume deterministic output, but a prompt-driven skill's behavior can vary run to run: a regular assertion either false-fails on harmless variation or gets loosened until it stops catching real regressions.

Decisions:
  • Kept deterministic assertions for everything that is actually deterministic (inputs, outputs, file writes) rather than routing every check through a model.
  • Used a model judge only for the subjective parts, and refused to trust it until it correctly grades known pass and fail examples. A judge that says PASS to everything agrees with pass-only examples, so both kinds are required.
  • Designed it to run as a normal red/green suite inside existing CI, so a broken skill fails the build the same way a broken function does.

Outcome: Prompt-based skills get a red/green suite instead of "looked fine when I tried it once", and a judge can't turn a check green unless it has shown it can tell good output from bad.

Python · source

AI Agents

Signal Scout

A Claude Code / OpenCode skill that turns a startup URL into an evidence-backed shortlist of first customers, market segments, and companies to pitch, using public signals only.

Problem: LLM-generated market research reads well but is routinely wrong: a model will confidently name a "prospect" or cite a "signal" that isn't actually there when you go look. For go-to-market research specifically, a fabricated lead costs real outreach time before anyone notices.

Decisions:
  • Built verify_sources.py as a hard gate: it re-fetches every cited source and confirms the claimed evidence is actually on the page before a claim is allowed into the report. Nothing ships unchecked.
  • Extracted the report rendering into a shared, deterministic component reused across the whole skill family, so formatting bugs and prompt drift get fixed once instead of per-skill.
  • Shipped as a standalone Claude Code / OpenCode plugin rather than a hosted service, so it runs against a user's own Claude Code session with no separate backend to operate.

Outcome: The verification step became the template for catching AI fabrication elsewhere: its containment-checking approach was later generalized into its own standalone tool, verify-before-ship.

Python, Claude Code Skills, OpenCode · source

First to First Sale

Takes a signal-scout prospect report and turns it into the right next action per prospect type: outreach sequences for individuals, content/GTM briefs for segments, BD pitches for companies. Never sends anything automatically.

  • Routes each prospect type (individual, segment, company) to the right output format
  • Drafts outreach sequences, content briefs, and BD pitches from real research, not templates
  • Human-in-the-loop by design: it only drafts, it never sends anything on its own
  • Companion skill to signal-scout, packaged as its own plugin (signal-outreach)

Impact: Closes the loop from "who should we talk to" to "here is the draft to send."

Python, Claude Code Skills · source

Signal Skills

Six Claude Code / Codex CLI skills that mine podcasts and app reviews for findings, track falsifiable predictions, and turn accumulated signal into prioritized, cited build proposals.

  • Mines podcasts (Apple, Spotify, RSS, YouTube) and app/store reviews into a queryable corpus
  • Tracks falsifiable predictions across episodes and scores guests on hit rate over time
  • A shared, SSOT-verified report renderer keeps all six plugins visually consistent

Impact: Replaces "mine this when I remember to ask" with a corpus that keeps compounding on its own.

Python, Claude Code Skills · source

Data Engineering

MetroPulse NYC

Segments every NYC subway station into behavioral archetypes using a serverless lakehouse: real ridership, geospatial, and demographic data, no synthetic inputs.

  • Serverless lakehouse pipeline: Dagster orchestration, DuckDB + Polars for transforms
  • FastAPI service layer over the resulting station archetypes
  • Fully geospatial: every station segmented on real MTA and neighborhood data

Impact: Handles the full pipeline from raw transit and geospatial data to station-level behavioral segments.

Python, Dagster, DuckDB, Polars, FastAPI · source

About

Shelfie is a vision-based iOS kitchen app, now in TestFlight beta. You walk the camera past your fridge and pantry once, and it builds the inventory that meal plans, grocery lists and expiry reminders run on.

Claude Code agents write most of its code, which turned a lot of my job into review. The harness they work in assumes an agent's report is unproven until a check backs it up, and a merge waits for me whenever a review asks for a person. Much of the CI targets work that looks finished without being finished, like a test filter that matches nothing and still exits green.

The projects here came out of that work: checking citations against their sources, testing prompt-based skills, keeping parallel agent sessions off the same file, and capping what an app spends on model calls.

Before this I spent five years as a senior data analyst at Reshet 13, an Israeli TV network, building the Python and SQL pipelines and forecasting models the business ran on.

Swift · SwiftUI · TypeScript · Python · SQL · PostgreSQL · Supabase · Claude Code · Gemini · Apple Vision · GitHub Actions · Playwright · Docker

Get in touch

Have a problem worth solving, or just want to talk shop about agent infrastructure? Email is the fastest way to reach me.

orenssegal@gmail.comgithub.com/OrenSegallinkedin.com/in/oren-segal