
Four Things That Broke: Shipping an Agent Tool on Public Data
I built a tool that matches a company against every public tender published in Germany — the agents were the easy part. The four failures between a working demo and a live one.
Hasan Halacli — AI Solution Architect & Technical Lead
10+ years turning AI transformation into production-grade systems — LLMs, MLOps, and cloud-native platforms built for security, reliability, and EU data residency.
I specialize in building end-to-end AI solutions, from data collection to deployment and maintenance, leading teams to deliver impactful results.
From classical ML to Deep Learning and Fine Tuning LLMs
Building scalable and efficient data pipelines.
Using cloud services to build scalable and efficient solutions.
Building efficient workflows with Power Platform and N8N.
I help engineering organisations make AI part of how they build software — deliberately, and in a way that survives an audit. My day job is running an AI platform for a critical-infrastructure operator under EU regulation, where an unreviewed change is a regulatory event rather than a bug. That is the standard I bring.
Your developers already use AI to write code. This makes it deliberate: what the tools may touch, who reviews the output, and how you prove it later.
The demo works. Now it needs real users, real data and someone on call. This is the gap where most GenAI projects quietly stall.
Find out which of your systems are in scope, what each one obliges you to do, and what is missing — before a customer or auditor asks first.
Point the Repo Reality Check at one of your public repositories. Five agents read it and hand you findings drawn from your own code — useful whether or not we ever speak.
Send me the report. I will tell you what I would fix first and what I would leave alone — including if the answer is that you need nobody.
Fixed scope with a deliverable your team keeps: the rules, the gates, the review protocol, running in your pipeline.
Find out in two minutes. The audit asks eight plain questions about how your teams work today and hands you a report written for people who do not read code.
Notes on shipping AI to production — architecture, LLMs, MLOps, and lessons from the field.

I built a tool that matches a company against every public tender published in Germany — the agents were the easy part. The four failures between a working demo and a live one.

You already know you should review agent code — I've written three posts on why. This is the how: what to read first, the failure signatures that look correct, and the tests you can actually…

A rule in a markdown file is a request. The agent honors it right up until 'temporarily' hardcoding the token becomes the path of least resistance. Requests need enforcement behind them.
Small, free, browser-based things I built from my work — no sign-up.
Paste your company website. Four agents read it, scan every public tender published in Germany in the last two weeks, and return the ones that fit — with the reason, the deadline and the link. Live data, no sign-up.
Find my tenders NewPaste any public GitHub repo and watch five agents read it live — review coverage, unpinned dependencies, missing CI gates. Real findings from real data, not a questionnaire.
Scan a repository NewEight plain questions about how your team uses AI to write code — scored report, biggest gaps, and a starter policy. Written for non-technical readers.
Take the audit InteractiveGenerate a drop-in rules file (AGENTS.md / CLAUDE.md) that turns a sycophantic coding agent into an expert. Toggle behaviors, then copy or download.
Open the builder ComplianceFour questions to place your AI system in the right EU AI Act risk tier — and see the obligations that follow. Guidance, not legal advice.
Classify a system InteractivePick stack, risk lane and change type — get a paste-ready review checklist for agent-written code.
Generate a checklist →Technical Lead / AI Solution Architect
Leading a team of 2 Senior Data Scientists, 1 Data Engineer, 1 MLOps Engineer, and 2 Master Students. Managing end-to-end AI solutions with strong focus on team development and strategic planning.
Lead Data Scientist / Tech Lead
Building end-to-end AI solutions, from data collection to deployment and maintenance. Working with NLP, Vision, Audio frameworks, RAG Chatbots and LLMs (including finetuning).
Senior Data Scientist, Full Stack Developer
NLP models (Hate Speech Detection using X data), showing result in an Web application and Google extension. Building receipt evaluation system using OCR.
Data Scientist/ Assistant Teacher
Teaching AI and ML to students and colleagues. Building AI solutions for the University.
Master of Science in Data Science
Duisburg Essen University of Applied Sciences
Electrical/Electronic Engineering
University of Applied Sciences
Years in AI/ML
Projects Delivered
Certifications
Engineers Led
Featured case studies
Broadcast live on national TV — real-time speech-to-speech, on air
A zero-retry, on-air AI interviewer: streaming STT → LLM dialogue management → low-latency TTS, with guardrails for live broadcast.
Read case study → Audio · Speech~40% word-error-rate reduction on internal test sets
LoRA fine-tuning + data augmentation to specialise Whisper for German dialects and domain speech — efficient, honestly measured, deployable.
Read case study → NLP · LLMGrounded, cited answers over enterprise documents
Hybrid retrieval + reranking + citations + web search + feedback loops, deployed under EU data constraints — answers people can actually trust.
Read case study → Agentic · LLMIntent-based routing to specialized agents
An orchestrator that classifies intent and routes complex queries to specialised agents — reliable, cost-bounded, and maintainable vs. a monolithic prompt.
Read case study →More work
Adapted open-source LLMs for domain-specific tasks, achieving significantly better accuracy than base models on internal benchmarks.
Orchestrated multiple AI agents for complex queries, routing requests to specialized models based on intent classification.
Enterprise chatbot with document upload, web search integration, and feedback loops for continuous answer improvement.
Multilingual topic clustering using BERTopic to organize massive media archives by semantic themes across languages.
NLP system detecting German holidays and cultural events in media content for seasonal programming recommendations.
Automatic language identification service classifying content across 50+ languages for localized recommendation engines.
Automated metadata extraction using KeyBERT to identify relevant keywords and phrases for content searchability.
AI negotiation assistant that extracts meeting insights and generates procurement arguments using RAG and hybrid search.
Customized Stable Diffusion model to generate brand-specific images matching company visual guidelines and style requirements.
Computer vision pipeline identifying main characters in video content and selecting optimal aesthetic frames for thumbnails.
Quality control tool analyzing video streams to detect and timestamp black frames for broadcast compliance checks.
Neural network scoring visual appeal of frames and clustering similar images for automated thumbnail selection.
Deep learning face analysis extracting demographics and emotions from video frames for content categorization metadata.
Enhanced Whisper ASR with German dialect training data, reducing word error rate by 40% on internal test sets.
Real-time subtitle generation system translating German broadcast content to English with 5-second latency using Whisper.
Audio analysis system detecting silence patterns to automatically segment videos into scenes for better content navigation.
Interactive AI interviewer using speech recognition and synthesis, broadcast live on national television.
AI-powered trading bot using sentiment analysis and fear index to automate buy/sell decisions in real-time.
Smart browser extension that provides contextual AI assistance by learning from expert users' knowledge and workflows.
Built intelligent automation workflows using Microsoft Power Platform and Copilot Studio for enterprise task orchestration.
Automated video ad creation from URLs using GPT-4 Vision for scene matching and voice cloning for narration.
Content moderation system using OCR and NLP to detect and flag adult material across multimedia datasets.
Age-appropriate content classifier identifying children's programming for parental control and recommendation filtering systems.
Master thesis building predictive churn models for manufacturing B2B clients using Azure ML and big data pipelines.
Mentored graduate students on NLP research including hate speech detection and sentiment analysis with real datasets.
End-to-end Kubernetes deployment on Azure AKS with GitOps via Flux and automated CI/CD pipelines using GitHub Actions.
Infrastructure as Code with Terraform for provisioning and managing cloud resources across environments with reproducible deployments.
I'm always interested in hearing about new projects, opportunities, and collaborations. Feel free to reach out!