AI Engineering
Five systems I built around language models, two for clients and three of my own. Each one shows how the data moves and the decisions behind it.
MarketAlerts · founding engineer · client work
Classifying financial signals
MarketAlerts was an investment-intelligence platform covering 20+ global markets. A data team ran the Python platform on Airflow and I owned its AI layer, which turned SEC filings, earnings-call transcripts, research and news into classified financial events. On the product side I built the alert engine that evaluated those signals for users. The company wound down in March 2026.
- Filings, transcripts and news
- Airflow pipelines
- First cut by a cheap model
- Hard cases to a stronger model
- Classified events
- Alert engine
Decisions
- Classification is hierarchical. A cheap model takes the wide first cut and a stronger one runs only where the decision is hard, and when a result comes out wrong there is a level to inspect instead of one long prompt to stare at.
- Providers sit behind one registry keyed by name, and each key resolves to that provider's whole config. Azure OpenAI needs an API version and a deployment name that OpenAI does not, so a flat shared interface would either hide that or leak it.
- Model output gets stripped of markdown fences before parsing. A model asked for JSON wraps it in a code block often enough to break a parser that trusts it.
- The earnings-call summariser turns raw transcripts into structured highlights. It started on OpenAI and moved to Gemini in production.
- For users, an analyst chat streams answers backed by live web search through OpenAI and Perplexity, with history shared across devices, voice questions and plan-based usage limits.
AI tutoring platform · senior software engineer · client work
Cost control and memory for an AI tutor
An exam-preparation platform with an AI tutor, used by 13M+ secondary-school students. The client is under NDA, so this stays at the level of the engineering problems.
- Student request
- Affordability check
- Model call
- Charge on success
- Conversation
- LLM pulls out lasting facts
- Stored memories
- Nightly merge and prune
Decisions
- Usage was counted in messages, which says nothing about cost on a product billed per token. I replaced it with cost-based metering on every plan tier, so the limit the product enforces is what each request costs.
- The affordability check runs before the model call and the charge only after it succeeds, so a failed request costs the student nothing.
- Memory runs as a background pipeline on Inngest. An LLM pass extracts lasting facts from conversations, and a nightly pass merges and prunes them, spread across a four-hour window so database load stays flat overnight.
- Credit top-ups sit alongside subscriptions through Stripe, including bank transfers that settle days after checkout, and the admin side breaks consumption down by user type.
Metis · my own platform
Retrieval over financial news
Metis takes in around 100,000 financial news articles a day from 5,000+ feeds, newsletters and archives and lines them up against market prices. Each article passes through separate processors for language, sentiment, entities, assets, topics, duplicates and embeddings, and the embeddings are what the retrieval runs on.
- Feeds and archives
- Language, entity and sentiment processors
- Self-hosted embedding service
- Qdrant
- Semantic search and agent tools
Decisions
- Embeddings go to Qdrant, next to PostgreSQL, ClickHouse and Meilisearch, which all kept their jobs when the vector index arrived.
- Natural-language search in the client API embeds the query into the same vector space as the stored articles, drops matches under a minimum similarity and puts one deadline on the embed and search calls together.
- A prediction-market strategy runs as an LLM agent with tools. Before it decides on a market it pulls related news from the last few days by embedding similarity and decides with those articles in its context.
- Topic tagging and asset matching each have an evaluator that runs candidate thresholds over a sample of real articles and reports what changes, down to precision and recall for the asset matcher.
- The models for embeddings, entity recognition, sentiment and language detection run as Python and FastAPI services, and clustering runs as a Rust service. LLM summaries go through OpenAI, Anthropic or Ollama.
Guestavo · my own product
An MCP server for running a venue
Guestavo is the software an independent restaurant, bar, cafe or event-led venue runs its guests on: reservations, menus, events, campaigns, loyalty and staff scheduling. Its MCP server lets an operator run menus, availability, bookings, events, contacts, shifts and reports from an AI client. Live in beta.
- AI client
- MCP server over stdio or HTTP
- Key scope and rate limit
- Venue data changed
- Audit entry with before and after
Decisions
- Every change an agent makes is logged with its before and after state, and it can be undone.
- API keys are scoped per organisation, property and action, rate-limited, and behind a kill switch.
- The server speaks both stdio and streamable HTTP, so it works from a desktop AI client as well as a hosted one.
My own tooling · how I build
The harness around coding agents
I have coded with AI every day for several years, on client work as much as my own: Tabnine first, then Cursor, then Claude Code, with Codex alongside it now. Most of what I built is the harness around the models, and Guestavo was built on it.
- Task
- Workflow with a gate and model tier per phase
- Agent with skills and MCP
- Deterministic checks
- Second-model review
- Human review
Decisions
- 180+ versioned workflow definitions across 10+ domains, each phase with its own gate and model tier.
- 30+ skill folders that an agent loads for the domain it is working in.
- MCP servers let agents read and write real project state, such as claiming a task and posting progress in my project manager, instead of guessing at it.
- Every change goes through deterministic rules first, then a second model, then me. Agent mistakes that keep coming back get promoted into lint rules.
- Nothing public ships without a human reviewing it.
Controls around the model calls
The same few controls come up across these systems. Each one links to where it runs.
Cost checked before the call
Requests are metered by what they cost in tokens, checked before the model runs and charged only on success.
Agent actions can be undone
Changes made over MCP are logged with their before and after state.
Scoped access for agents
API keys are limited per organisation, property and action, rate-limited, and behind a kill switch.
Providers can be swapped
Models sit behind a registry that keeps each provider's config intact, and one summariser moved from OpenAI to Gemini in production.
Thresholds are measured
Candidate classifier thresholds run against samples of real articles, with precision and recall where there is something to compare against.
Three layers of review
Deterministic rules, then a second model, then a human, on every change an agent writes.
Writing about it
- 🗂️ My Agent Skill Tree
A tour of my coding agent skills and how they fit together.
- 📰 Metis Revisited: The Platform Behind the News
An update on Metis and its financial data tools.
- 🧠 Five Attempts at Giving an Assistant a Memory
My five attempts at building a personal assistant that remembers.
- 📄 Turning SEC Filings Into Something You Can Query
How I turn SEC filings into queryable data in Metis.
- 🤖 How I Work With Coding Agents
The skills and workflows I use with coding agents.
Want this on your team?
Hiring an AI engineer for a long-term role, employed or on a B2B contract? Send me the role description.
Discuss a role