Ömer Faruk Koç

MLOps & AI Platform Engineer

Building production AI/ML platforms, agentic systems, retrieval infrastructure and reliable data systems.

3+ years building and operating production ML, data and GenAI systems, with current work focused on agent reliability, evaluation, observability and platform control.

Türkiye · Open to remote opportunities · MLOps · AI Platform · GenAI

Request lifecycle

  1. Request
  2. AI Platform API
  3. Agent Runtime
  4. RAG Retrieval
  5. Model Serving
  6. Guardrail
  7. allow
  8. confirm
  9. human
  10. deny
  11. Response

Trace + Metrics

Career highlights

At a glance.

Experience3+ YEARS
Production AI/MLProfessional experience across ML, data and GenAI systems.
Performance75% RUNTIME REDUCTION
Core data workflowSequential Oracle processing was re-architected into modular parallel dbt workflows, reducing end-to-end runtime by 75%.
ProductionEND-TO-END ML
Lifecycle ownershipOwned validation, versioning, promotion, serving, retraining and monitoring across multiple production ML use cases.
InfrastructurePRIVATE GENAI
On-prem GPU servingAdapted validated cloud prototypes to private GPU infrastructure for quantized open-source model serving under data-residency constraints.

Selected work

Systems, not demos.

Public engineering projects built around control boundaries, reliability, evaluation, observability and explicit trade-offs.

Project / 01current

AI Infrastructure / Platform Engineering

ML Platform Infrastructure

Local ML platform reference implementation on Kubernetes: an inference service, its full MLflow/PostgreSQL/MinIO lifecycle, GitOps with Argo CD, HPA autoscaling, security hardening and observability, validated with real failure drills instead of documentation claims.

  • Kubernetes
  • kind
  • Helm
  • Argo CD
  • Terraform

Reproducibility

PROVEN, NOT ASSUMED

Destroyed cluster · images · build cache

make local-up followed by make local-test reached 11/11 acceptance checks in 901s (~15 min) after the cluster, images and build cache were destroyed first. Status is local-v1.0.0: the Kubernetes implementation (M0–M12) is validated and frozen. AWS (M13+) has not started — no cloud resource has been created, and Terraform is at the design/static-validation level only (fmt, validate, tflint; no plan, no apply).

Project / 02current

AI Reliability / Execution Infrastructure

Agentic Customer Service Platform

Customer-service agent platform where the LLM proposes refunds, cancellations, lookups, tickets and escalations while deterministic software owns scope, policy, confirmation, revalidation, idempotency and execution.

  • Python
  • FastAPI
  • LangGraph
  • SQLAlchemy
  • PostgreSQL

Execution authority

SERVER-OWNED

LLM proposal → deterministic execution

Authentication, scope, policy, confirmation, revalidation, idempotency and business execution stay outside the model. The exercised D2c slice recorded 0 unauthorized mutations and 0 unsafe executions.

Project / 03current

Agent Systems / AI Platform

Agentic SRE

Evidence-driven root-cause analysis engine for Kubernetes incidents: a bounded, read-only investigator gathers evidence over changes, events, logs, traces, dependencies and topology, while a deterministic RCA engine — not the model — makes the final root-cause judgment.

  • Python
  • FastAPI
  • LangGraph
  • Pydantic
  • SQLAlchemy

Root-cause judgment

DETERMINISTIC

Investigator proposes reads · deterministic RCA decides

A bounded, read-only investigator acquires evidence one legal read at a time; deterministic normalization, hypothesis rebuilding, verification and root-cause selection own the diagnosis. On the frozen blind ITBench-Lite TEST25 holdout it reached 21/25 (84%) exact-root agreement with 0 model calls. This is a measured 25-scenario benchmark result, not a universal production accuracy guarantee.

Project / 04current

Generative AI / RAG Platform

Knowledge Base RAG

Local-first multilingual RAG platform where tenant scope, evidence construction, support-unit identity and occurrence-aware validation are separate boundaries, and where a change is adopted only if it passes a decision rule frozen before the result was known.

  • Python
  • FastAPI
  • React
  • TypeScript
  • Vite

Multilingual reranker benchmark

MRR 0.367 → 0.956

220 questions · frozen multilingual benchmark

The multilingual reranker improved Cross MRR from the prior English reranker's 0.367 to 0.956 and reached Recall@5 of 1.000 on both TR-to-EN and EN-to-TR slices. These are retrieval/ranking metrics, not final-answer accuracy.

Project / 05current

MLOps / AI Platform

ModelOps Control Plane

Policy-driven ML release control plane combining progressive canary delivery, delayed-ground-truth quality gates, automated promotion and rollback, and desired-vs-observed routing reconciliation.

  • Python
  • FastAPI
  • SQLAlchemy
  • SQLite
  • Alembic

Routing control loop

DESIRED ↔ OBSERVED

Durable database state → router reconciliation

The database owns desired traffic; the router is restart-losable observed state. A worker-triggered reconcile tick repairs drift after a restart or failed push.

Project / 06current

Structured Data / Governed Text-to-SQL

DecisionSQL

Governed one-shot Text-to-SQL for enterprise analytics: the model emits one typed decision, deterministic software owns SQL admission and read-only execution, and correctness is measured by execution against counterfactual database states.

  • Python
  • FastAPI
  • PostgreSQL
  • SQLAlchemy
  • Alembic

Execution authority

SERVER-OWNED

Model proposes SQL · deterministic admission decides

An ANSWER + SQL submission crosses parse, global policy, request-scoped relation authority, grain safety, EXPLAIN and a cost gate, and only an accepted immutable QueryPlan reaches the restricted read-only executor. There is no raw unsafe fallback, and the executor never runs SQL the safety service did not admit.

Project / 07current

Distributed Systems / Streaming

Real-Time Commerce Platform

Production-oriented event-driven commerce platform where Kafka may redeliver, but layered idempotency, transactional persistence and bounded failure handling protect durable business effects.

  • Python
  • TypeScript
  • FastAPI
  • Next.js
  • Kafka

Sustainable isolated capacity

Workload-dependent

~1,050 evt/s at 42.8% fraud-eligible · ~1,600 at 0%

The processor has no single workload-independent ceiling, because fraud-eligible events take a longer path. The ~750 to ~1,050 evt/s (+40%) optimisation was measured on the 42.8% workload. Local isolated benchmark, not production capacity.

Engineering areas

Platform work across five connected domains.

Capabilities are grouped by engineering evidence—not percentages or keyword counts.

01

AI / ML Platform

Model lifecycle, progressive delivery, delayed quality feedback, policy-driven release control and observable production operations.

  • FastAPI
  • Docker
  • Kubernetes
  • MLRun
  • GitHub Actions
02

Generative AI / RAG

Retrieval, reranking, citation integrity, evaluation and private open-source model serving.

  • Qdrant
  • Ollama
  • OpenTelemetry
  • DeepEval
  • LangChain
03

Distributed Systems

Event processing with explicit delivery guarantees, failure handling and measured service limits.

  • Kafka
  • PostgreSQL
  • Redis
  • Prometheus
  • Grafana
04

Data Engineering

Batch and near-real-time pipelines, transformation systems, quality controls and data lineage.

  • dbt
  • Airflow
  • Oracle
  • sqlglot
  • NetworkX
05

Agent Systems / Agent Infrastructure

Stateful agent workflows with deterministic control boundaries, confirmation and recovery, secure tool interfaces, evaluation and observability.

  • LangGraph
  • MCP
  • FastMCP
  • PostgreSQL
  • OpenTelemetry

Professional experience

Production systems at scale.

Fibabanka · Analytics Center of Excellence

MLOps & Analytics Engineer

Built and operated production ML, data and Generative AI platform capabilities across analytics workloads.

Period
Mar 2023 — Mar 2026
Location
Istanbul, Türkiye

Professional experience informs the engineering questions explored in public projects; the public repositories are independent work.

Explore experience

Engineering writing

Notes from building the systems.

Short technical pieces on reliability, evaluation, retrieval, distributed systems and production AI.

Agent Reliability

Hard Gates + Frozen Hashes for AI Coding Agents

Making agentic development verifiable rather than merely autonomous: external gates decide whether work is acceptable, and hashed baselines stop the agent from redefining what acceptable means.

  • AI Agents
  • Guardrails
  • Evaluation
Read
Agent Reliability

Designing Guardrails for Production AI Agents

A practical execution model for tool-using agents built from typed proposals, deterministic policy, durable confirmation, revalidation, idempotency, and audit.

  • AI Agents
  • Guardrails
  • Reliability
Read
Retrieval & RAG

63 Rescues, 0 Drops, and 2.4 Seconds

A cross-lingual reranker moved Recall@5 from 0.9563 to 1.0000. The aggregate is the least interesting number in that sentence, and the latency figure means less than it looks like.

  • RAG
  • Evaluation
  • Performance Engineering
Read
Explore writing

Engineering graph

See how the work connects.

390 source-grounded nodes and 519 validated relationships across experience, projects, technologies, concepts, evidence and current learning directions.

Explore Engineering Graph

Contact

Have an interesting platform, data or AI infrastructure problem?

Open to remote opportunities · UTC+3 / Istanbul

Available for new opportunities