My engagements can be hands-on, or purely guidance and consulting. Examples of my earlier engagements: bringing large projects over the finish line, setting up telemetry, setting up alerts and on-call rotation, and working hand-in-hand with your engineers during incident resolution. You can contact me for an analysis of your system and get written findings that you can act on, whether or not we keep working together.

AI-assisted engineering

LLM coding tools can be great with proper use, but unreliable in real large codebases. How good a model is depends on the tooling, the context, and the verification around it. I work on making AI and LLM productive in an engineering organization: context engineering, setups with large mono-repos or multiple focused repos, preparing and maintaining repository guidance files, AI review systems, agent workflows with verification built in, and measurement of actual throughput improvements. At Caffeine.ai I built the custom LLM PR-review system that became the team’s main review system, and an AI-native ticket-to-PR loop (integration with Linear). The loop covers everything between plan, implement, self-review, fix, update ticket, and respond to reviewer comments. My setups can achieve sustained 5–15 meaningful PRs per day per engineer while burning down tech debt.

What you get: repository analysis, guidance, and context engineering fitted to your codebase; review automation tailored to your standards; agent workflows for the work worth automating; and numbers that show whether any of it is paying off.

Production reliability

Production systems fail in ways that are obvious in hindsight and missing from the dashboards, until someone fixes the dashboards. I work on observability, incident reduction, rollout safety, migrations without downtime, and on-call and operations maturity. At DFINITY I led reliability for the Internet Computer, a 1,400+ node HA network: critical incidents went from double digits per year to 1–2 while usage grew by orders of magnitude and we shipped hundreds of network upgrades per year. I built observability across a multi-AZ Kubernetes estate, and with two colleagues rewrote the CI system, cutting merge times from 1–2 days to 30–60 minutes.

What you get: working telemetry with alerts people trust, a measured drop in incident count and severity, rollouts you can reverse, and migrations that do not need a maintenance window.

Storage and distributed systems

Storage is where systems succeed or fail: latency budgets, durability, failure modes, and the trade-offs that only show up under real load. I do architecture reviews and design work for storage engines, data platforms, and the distributed systems underneath them, from flash firmware behavior up to cluster-level design. At IBM Research I was one of the core FTL firmware developers for the IBM FlashCore Module. Later I architected an analytics platform ingesting real-time telemetry from 11,000+ enterprise storage systems. That grounding comes from a research career: 16+ peer-reviewed papers and 60+ patents.

What you get: an architecture review that covers the whole stack, written findings with concrete recommendations, and a design your engineers can implement.

Working together

Engagements range from a short assessment with written findings to ongoing advisory. If any of this sounds like your problem, find me on LinkedIn.