AI Engineering

Five systems I built around language models, two for clients and three of my own. Each one shows how the data moves and the decisions behind it.

MarketAlerts · founding engineer · client work

Classifying financial signals

MarketAlerts was an investment-intelligence platform covering 20+ global markets. A data team ran the Python platform on Airflow and I owned its AI layer, which turned SEC filings, earnings-call transcripts, research and news into classified financial events. On the product side I built the alert engine that evaluated those signals for users. The company wound down in March 2026.

How the data moves
  1. Filings, transcripts and news
  2. Airflow pipelines
  3. First cut by a cheap model
  4. Hard cases to a stronger model
  5. Classified events
  6. Alert engine

Decisions

  • Classification is hierarchical. A cheap model takes the wide first cut and a stronger one runs only where the decision is hard, and when a result comes out wrong there is a level to inspect instead of one long prompt to stare at.
  • Providers sit behind one registry keyed by name, and each key resolves to that provider's whole config. Azure OpenAI needs an API version and a deployment name that OpenAI does not, so a flat shared interface would either hide that or leak it.
  • Model output gets stripped of markdown fences before parsing. A model asked for JSON wraps it in a code block often enough to break a parser that trusts it.
  • The earnings-call summariser turns raw transcripts into structured highlights. It started on OpenAI and moved to Gemini in production.
  • For users, an analyst chat streams answers backed by live web search through OpenAI and Perplexity, with history shared across devices, voice questions and plan-based usage limits.
PythonAirflowKubernetesOpenAIAzure OpenAIGeminiPerplexityTypeScriptBullMQ

AI tutoring platform · senior software engineer · client work

Cost control and memory for an AI tutor

An exam-preparation platform with an AI tutor, used by 13M+ secondary-school students. The client is under NDA, so this stays at the level of the engineering problems.

Every tutor request
  1. Student request
  2. Affordability check
  3. Model call
  4. Charge on success
Memory, in the background
  1. Conversation
  2. LLM pulls out lasting facts
  3. Stored memories
  4. Nightly merge and prune

Decisions

  • Usage was counted in messages, which says nothing about cost on a product billed per token. I replaced it with cost-based metering on every plan tier, so the limit the product enforces is what each request costs.
  • The affordability check runs before the model call and the charge only after it succeeds, so a failed request costs the student nothing.
  • Memory runs as a background pipeline on Inngest. An LLM pass extracts lasting facts from conversations, and a nightly pass merges and prunes them, spread across a four-hour window so database load stays flat overnight.
  • Credit top-ups sit alongside subscriptions through Stripe, including bank transfers that settle days after checkout, and the admin side breaks consumption down by user type.
TypeScriptNext.jstRPCPostgreSQLDrizzle ORMInngestStripe

Metis · my own platform

Retrieval over financial news

Metis takes in around 100,000 financial news articles a day from 5,000+ feeds, newsletters and archives and lines them up against market prices. Each article passes through separate processors for language, sentiment, entities, assets, topics, duplicates and embeddings, and the embeddings are what the retrieval runs on.

How the data moves
  1. Feeds and archives
  2. Language, entity and sentiment processors
  3. Self-hosted embedding service
  4. Qdrant
  5. Semantic search and agent tools

Decisions

  • Embeddings go to Qdrant, next to PostgreSQL, ClickHouse and Meilisearch, which all kept their jobs when the vector index arrived.
  • Natural-language search in the client API embeds the query into the same vector space as the stored articles, drops matches under a minimum similarity and puts one deadline on the embed and search calls together.
  • A prediction-market strategy runs as an LLM agent with tools. Before it decides on a market it pulls related news from the last few days by embedding similarity and decides with those articles in its context.
  • Topic tagging and asset matching each have an evaluator that runs candidate thresholds over a sample of real articles and reports what changes, down to precision and recall for the asset matcher.
  • The models for embeddings, entity recognition, sentiment and language detection run as Python and FastAPI services, and clustering runs as a Rust service. LLM summaries go through OpenAI, Anthropic or Ollama.
TypeScriptNestJSPythonFastAPIRustQdrantPostgreSQLClickHouseMeilisearchBullMQRabbitMQAnthropic

Guestavo · my own product

An MCP server for running a venue

Guestavo is the software an independent restaurant, bar, cafe or event-led venue runs its guests on: reservations, menus, events, campaigns, loyalty and staff scheduling. Its MCP server lets an operator run menus, availability, bookings, events, contacts, shifts and reports from an AI client. Live in beta.

How the data moves
  1. AI client
  2. MCP server over stdio or HTTP
  3. Key scope and rate limit
  4. Venue data changed
  5. Audit entry with before and after

Decisions

  • Every change an agent makes is logged with its before and after state, and it can be undone.
  • API keys are scoped per organisation, property and action, rate-limited, and behind a kill switch.
  • The server speaks both stdio and streamable HTTP, so it works from a desktop AI client as well as a hosted one.
TypeScriptMCPFastifyNext.jsPostgreSQLDrizzle ORMBullMQ

My own tooling · how I build

The harness around coding agents

I have coded with AI every day for several years, on client work as much as my own: Tabnine first, then Cursor, then Claude Code, with Codex alongside it now. Most of what I built is the harness around the models, and Guestavo was built on it.

How the data moves
  1. Task
  2. Workflow with a gate and model tier per phase
  3. Agent with skills and MCP
  4. Deterministic checks
  5. Second-model review
  6. Human review

Decisions

  • 180+ versioned workflow definitions across 10+ domains, each phase with its own gate and model tier.
  • 30+ skill folders that an agent loads for the domain it is working in.
  • MCP servers let agents read and write real project state, such as claiming a task and posting progress in my project manager, instead of guessing at it.
  • Every change goes through deterministic rules first, then a second model, then me. Agent mistakes that keep coming back get promoted into lint rules.
  • Nothing public ships without a human reviewing it.
Claude CodeCodexMCPTypeScriptESLint

Controls around the model calls

The same few controls come up across these systems. Each one links to where it runs.

Writing about it

Want this on your team?

Hiring an AI engineer for a long-term role, employed or on a B2B contract? Send me the role description.

Discuss a role
Get in touch