For more than 20 years, I've designed and built software systems for the automotive and industrial sectors, where reliability, maintainability, and production-grade engineering are non-negotiable. Today, I apply the same engineering discipline to AI systems.
I design and build production-grade Large Language Model (LLM) systems by combining software architecture principles with modern AI engineering.
Rather than treating LLMs as standalone applications, I design systems where language models operate as components within a larger architecture that includes orchestration, retrieval, inference, memory, and tool execution.
My work spans the complete lifecycle of LLM-based systems, including local model deployment, Retrieval-Augmented Generation (RAG), parameter-efficient fine-tuning (PEFT/LoRA), inference optimization, and Agentic AI systems.
Core areas of expertise include:
I focus on designing systems that are reliable, observable, and maintainable, while addressing real-world constraints such as latency, memory utilization, scalability, and deployment complexity.
My objective is to build AI systems that are structured, extensible, and engineered for production — not demonstrations driven solely by prompts.
I move across the full chain without gaps: problem formulation → architecture → implementation → evaluation → deployment. Every concept is driven by system necessity, not curriculum sequence — I understand transformers as systems that continuously reposition contextual representations for semantic alignment, and I know the path from this understanding to production systems.
Fluent in English, German, and Romanian.