Ömer Faruk Koç

MLOps & AI Platform Engineer

Building production AI/ML platforms, agentic systems, retrieval infrastructure and reliable data systems.

3+ years building and operating production ML, data and GenAI systems, with current work focused on agent reliability, evaluation, observability and platform control.

Türkiye · Open to remote opportunities · MLOps · AI Platform · GenAI

Request lifecycle

  1. Request
  2. AI Platform API
  3. Agent Runtime
  4. RAG Retrieval
  5. Model Serving
  6. Guardrail
  7. allow
  8. confirm
  9. human
  10. deny
  11. Response

Trace + Metrics

Career highlights

At a glance.

Experience3+ YEARS
Production AI/MLProfessional experience across ML, data and GenAI systems.
Performance75% RUNTIME REDUCTION
Core data workflowSequential Oracle processing was re-architected into modular parallel dbt workflows, reducing end-to-end runtime by 75%.
ProductionEND-TO-END ML
Lifecycle ownershipOwned validation, versioning, promotion, serving, retraining and monitoring across multiple production ML use cases.
InfrastructurePRIVATE GENAI
On-prem GPU servingAdapted validated cloud prototypes to private GPU infrastructure for quantized open-source model serving under data-residency constraints.

Selected work

Systems, not demos.

Public engineering projects built around control boundaries, reliability, evaluation, observability and explicit trade-offs.

Project / 01current

AI Reliability / Execution Infrastructure

Agentic Customer Service Platform

Customer-service agent platform where the LLM proposes refunds, cancellations, lookups, tickets and escalations while deterministic software owns scope, policy, confirmation, revalidation, idempotency and execution.

  • Python
  • FastAPI
  • LangGraph
  • SQLAlchemy
  • PostgreSQL

Execution authority

SERVER-OWNED

LLM proposal → deterministic execution

Authentication, scope, policy, confirmation, revalidation, idempotency and business execution stay outside the model. The exercised D2c slice recorded 0 unauthorized mutations and 0 unsafe executions.

Project / 02current

Generative AI / RAG Platform

Knowledge Base RAG

Local-first multilingual RAG platform where tenant scope, evidence construction, support-unit identity and occurrence-aware validation are separate boundaries, and where a change is adopted only if it passes a decision rule frozen before the result was known.

  • Python
  • FastAPI
  • React
  • TypeScript
  • Vite

Multilingual reranker benchmark

MRR 0.367 → 0.956

220 questions · frozen multilingual benchmark

The multilingual reranker improved Cross MRR from the prior English reranker's 0.367 to 0.956 and reached Recall@5 of 1.000 on both TR-to-EN and EN-to-TR slices. These are retrieval/ranking metrics, not final-answer accuracy.

Project / 03current

MLOps / AI Platform

ModelOps Control Plane

Policy-driven ML release control plane combining progressive canary delivery, delayed-ground-truth quality gates, automated promotion and rollback, and desired-vs-observed routing reconciliation.

  • Python
  • FastAPI
  • SQLAlchemy
  • SQLite
  • Alembic

Routing control loop

DESIRED ↔ OBSERVED

Durable database state → router reconciliation

The database owns desired traffic; the router is restart-losable observed state. A worker-triggered reconcile tick repairs drift after a restart or failed push.

Project / 04current

Distributed Systems / Streaming

Real-Time Commerce Platform

Production-oriented event-driven commerce platform where Kafka may redeliver, but layered idempotency, transactional persistence and bounded failure handling protect durable business effects.

  • Python
  • TypeScript
  • FastAPI
  • Next.js
  • Kafka

Sustainable isolated capacity

Workload-dependent

~1,050 evt/s at 42.8% fraud-eligible · ~1,600 at 0%

The processor has no single workload-independent ceiling, because fraud-eligible events take a longer path. The ~750 to ~1,050 evt/s (+40%) optimisation was measured on the 42.8% workload. Local isolated benchmark, not production capacity.

Engineering areas

Platform work across five connected domains.

Capabilities are grouped by engineering evidence—not percentages or keyword counts.

01

AI / ML Platform

Model lifecycle, progressive delivery, delayed quality feedback, policy-driven release control and observable production operations.

  • FastAPI
  • Docker
  • Kubernetes
  • MLRun
  • GitHub Actions
02

Generative AI / RAG

Retrieval, reranking, citation integrity, evaluation and private open-source model serving.

  • Qdrant
  • Ollama
  • OpenTelemetry
  • DeepEval
  • LangChain
03

Distributed Systems

Event processing with explicit delivery guarantees, failure handling and measured service limits.

  • Kafka
  • PostgreSQL
  • Redis
  • Prometheus
  • Grafana
04

Data Engineering

Batch and near-real-time pipelines, transformation systems, quality controls and data lineage.

  • dbt
  • Airflow
  • Oracle
  • sqlglot
  • NetworkX
05

Agent Systems / Agent Infrastructure

Stateful agent workflows with deterministic control boundaries, confirmation and recovery, secure tool interfaces, evaluation and observability.

  • LangGraph
  • MCP
  • FastMCP
  • PostgreSQL
  • OpenTelemetry

Professional experience

Production systems at scale.

Fibabanka · Analytics Center of Excellence

MLOps & Analytics Engineer

Built and operated production ML, data and Generative AI platform capabilities across analytics workloads.

Period
Mar 2023 — Mar 2026
Location
Istanbul, Türkiye

Professional experience informs the engineering questions explored in public projects; the public repositories are independent work.

Explore experience

Engineering writing

Notes from building the systems.

Short technical pieces on reliability, evaluation, retrieval, distributed systems and production AI.

Agent Reliability

Hard Gates + Frozen Hashes for AI Coding Agents

Making agentic development verifiable rather than merely autonomous: external gates decide whether work is acceptable, and hashed baselines stop the agent from redefining what acceptable means.

  • AI Agents
  • Guardrails
  • Evaluation
Read
Agent Reliability

Designing Guardrails for Production AI Agents

A practical execution model for tool-using agents built from typed proposals, deterministic policy, durable confirmation, revalidation, idempotency, and audit.

  • AI Agents
  • Guardrails
  • Reliability
Read
Retrieval & RAG

63 Rescues, 0 Drops, and 2.4 Seconds

A cross-lingual reranker moved Recall@5 from 0.9563 to 1.0000. The aggregate is the least interesting number in that sentence, and the latency figure means less than it looks like.

  • RAG
  • Evaluation
  • Performance Engineering
Read
Explore writing

Engineering graph

See how the work connects.

270 source-grounded nodes and 349 validated relationships across experience, projects, technologies, concepts, evidence and current learning directions.

Explore Engineering Graph

Contact

Have an interesting platform, data or AI infrastructure problem?

Open to remote opportunities · UTC+3 / Istanbul

Available for new opportunities