agentic systems architect
I make AI agents reliable enough to run in production.
I'm Vitalii Serbyn — I work with US, UK, and EU teams to get agent systems production-ready. Two specialties: agent reliability (trajectory evals, false-green-rate metering, evidence-binding gates, trust-tiered execution) and LLM cost engineering (multi-model routing and cost tracking) — bringing spend down without lowering output quality.
Updated July 2026
why this mixture is rare
Reliability and cost, held by the same hands
Agent reliability engineering
Trajectory evals, false-green-rate metering, evidence-binding gates, and trust-tiered L0–L4 execution — the machinery that turns "seems fine" into a number you can gate on.
LLM cost engineering
Multi-model routing, semantic caching, and per-service cost tracking that cut the bill without lowering output quality — the same levers whether you spend hundreds or hundreds of thousands.
12+ years of production systems
Long before agents: shipping software people depend on — including a consumer app with 100M+ downloads and a Google Play Editor’s Choice award, and an 80–90% RPC cost reduction on a Web3 platform, both as an engineer in employed roles.
quantified results
The numbers, each with its context
Every figure below says where it came from — a system I built, my own platform, or an employed role. No borrowed client metrics.
01_systems_i_designed_and_built
Systems I designed and built
Not client logos — my own systems, where I proved the reliability and cost methods I bring to your agents.
Ascend
Multi-agent orchestration daemon: trust-tiered L0–L4 execution, policy gates, budget guards, and audit logging across live projects.
DW
Autonomous spec-driven development pipeline: a spec goes in, working change comes out, gated by acceptance checks before it lands.
Crest
LangGraph-based content engine: multi-stage pipeline with multi-model routing and multi-platform publishing.
Built with Python, LangGraph, FastAPI, PostgreSQL, Redis, and Docker.
02_ways_to_work_together
Fixed-scope offers
AI Agent Production-Readiness Audit
A ground-truth read on whether your agents are safe to run in production: trajectory evals, false-green metering, evidence-binding gaps, and a prioritized fix list.
Learn moreLLM Cost Teardown
A line-by-line teardown of where your LLM spend goes and the six levers that bring it down — without lowering output quality.
Learn moreFractional AI Architect
Ongoing architectural ownership of your agent platform: reliability, cost, and the gated rollout discipline to keep both in line as you scale.
Learn moreAlso available: Advisory at $200/hr (10-hour minimum block). Audit fee credited against a retainer started within 60 days.
03_operator_protocol
How I work
Same sequence every engagement — measure before touching, gate every change on evidence.
- 01
Ground-truth audit first
Before changing anything, I measure what your agents actually do — real trajectories, real pass/fail, real cost — not what the dashboard claims.
- 02
Evals before changes
I build the trajectory and false-green evals that turn "seems fine" into a number, so every later change is judged against a fixed bar.
- 03
Evidence-binding
Every "green" has to point at the evidence that earned it. Claims that cannot cite their proof are treated as failures.
- 04
Gated rollout
Changes ship behind trust tiers and budget guards, promoting only when the evals and cost metrics hold — with a fast path back.
04_proof_of_method
See the method, not just claims
Sanitized sample audit report
The exact deliverable format — findings, severity, evidence, and a prioritized fix list — with a real (client-anonymized) example.
Available on requestMethod walkthrough (2–3 min)
A short screen recording walking through how an agent trajectory is evaluated and where the false-greens hide.
Coming soon05_what_colleagues_say
What colleagues say
Named recommendations from engineers and leads I've worked with are shared on request, with links to the originals.
Reference available on request.
LinkedIn recommendationReference available on request.
LinkedIn recommendationReference available on request.
LinkedIn recommendation06_continuity_and_trust
How engagements are run
- Contracting entity
- Easelect LTD (UK LTD). Invoiced in USD/GBP/EUR; standard contractor terms.
- IP & ownership
- Full IP transfer on payment. Your code and findings are yours.
- Access
- Read-only access by default. Write access only when a change is agreed, scoped, and gated.
- NDA / DPA
- NDA on request; DPA available for engagements touching personal data.
- Continuity
- Based in Kyiv with backup power and connectivity in place to keep engagements on schedule.
- References
- Client references handled privately, shared on request with permission.
07_writing
Notes on agent reliability
The Arc Since the Six Were Caught
After fixing six agent failure modes, the system graduated from outcome checking to path checking. A field report on trajectory evals and generalization.
The Calibration Ledger: 58 Runs, 93% Pass, n Is Small
58 runs, ~93% autonomous pass, and why that number is honest evidence — not a reliability proof. A field report on agent evals at solo scale.
False Greens: Three Structural Observations
How fail-closed design, headless execution seams, and converting incidents to fixtures stop AI agents lying about success. A field report from a solo dev-loop.
One-Sided Contracts Break Agent Pipelines
A strict contract enforced on only one side produces false rejections as reliably as a loose one produces false acceptances. A field report from my fleet.
Get a straight read on your agents
A 30-minute systems call: tell me what your agents do and where they worry you, and I'll tell you whether an audit, a cost teardown, or a retainer is the right next step.
Fixed-scope audits · Read-only by default · NDA on request · US/UK/EU remote