LLM Evaluation and Monitoring Platform by Abdulqoyyum AileruLLM Evaluation and Monitoring Platform by Abdulqoyyum Aileru

LLM Evaluation and Monitoring Platform

Abdulqoyyum Aileru

Abdulqoyyum Aileru

Celsius: Making AI Models Measurable, Observable, and Trustworthy at Scale

Client: SilverBrain AI AG Role: Back-End Engineer
The Problem
Deploying an AI model is only the beginning. Once a model is live, teams need to know whether it's still giving good answers, where it fails, and why. Models drift, outputs degrade quietly, and a single bad response can damage user trust before anyone notices.
Most teams fly blind here. Inference logs pile up in volumes no person can review, and when something breaks, engineers waste hours tracing the problem through scattered data. SilverBrain AI needed a platform that could evaluate and monitor AI models continuously, turning a flood of raw logs into clear, actionable insight.
The Vision
Celsius was built to be the control room for AI in production. Every inference gets captured, evaluated, and made searchable, so teams can measure model quality over time, spot problems early, and debug with evidence instead of guesswork.
My Role
I was a core back-end engineer on Celsius, focused on the data and infrastructure that make evaluation possible at scale: ingesting and storing inference logs, keeping retrieval fast, improving automated classification, and building the monitoring pipelines that keep the system observable.
The Solution
High-volume log ingestion. Celsius tracks over 500,000 AI inference logs every month. I worked on the back-end services that capture this data reliably, so every model interaction is recorded and ready for evaluation without slowing down the systems being monitored.
Storage designed for speed. Evaluation is only useful if teams can query the data quickly. I designed a storage layer pairing PostgreSQL for durable, structured records with Redis for caching frequently accessed data. This combination cut retrieval times by 50%, so engineers get answers in moments rather than waiting on slow queries.
Smarter automated classification. Reviewing half a million logs by hand is impossible, so Celsius relies on automated classification to sort and flag outputs. I applied LLM fine-tuning to the classification models, improving accuracy by 25%. Teams could trust the flags they received and focus on real issues instead of false alarms.
Real-time monitoring and observability. I built monitoring pipelines with Prometheus and Grafana, giving teams live dashboards on model behavior and system health. When something goes wrong, the data needed to diagnose it is already visible, reducing debugging time by 40%.
The Outcome
Celsius gave SilverBrain AI a clear, real-time view of how its AI models perform in production. With 500K+ monthly logs captured and evaluated, faster data retrieval, more accurate classification, and live monitoring, teams moved from reacting to problems to catching them early.
What This Shows
I build the infrastructure that makes AI dependable: scalable data pipelines, fast storage, careful model improvement, and observability that shortens the path from problem to fix. I understand that shipping AI responsibly means being able to measure it.
Like this project

Posted Feb 25, 2025

Developed backend for Celsius, an LLM evaluation & monitoring platform. Optimized API performance, enhanced data processing, and ensured scalable architecture.