I'm a third year AI student at Howest in Kortrijk, Belgium, originally from Baku. I spend most of my time on anomaly detection and model evaluation, and I don't trust a number until I have reproduced it myself. Most of what is on this page is either a known system rebuilt from the paper and measured on my own laptop, or an evaluation where the obvious story turned out to be wrong.
The numbers below all come from runs in the linked repos, and most of those repos fail their own CI if the README stops matching the results. Outside the from-scratch work I built a multilingual hallucination detector that runs on a Raspberry Pi, a benchmark for Azerbaijani LLMs, and offline accessibility tools for deaf and blind users. Before Belgium I taught programming and AI to 200+ students in Baku, and I run the IT side of a digital signage company with 1,000+ edge devices.
Beyond work, I like mathematics olympiads (a gold medal in 2022), events (I founded two in Baku) and cycling around Kortrijk.
ARC Prize 2026, ARC-AGI-2. My own DSL program search scores 0.00 on the hidden set, and the paper track write-up explains why. On the leaderboard I run the public NVARC notebook as a baseline. My best is 31.39, but identical reruns land anywhere from 27 to 31, so one good score doesn't mean much. I also tried training it more on the 2,000 puzzles it gets most wrong. It got better at those puzzles, but on 30 test tasks it solved exactly as many as before, so I didn't submit it.
Two papers in review at SSRN. An evaluation of my PatchCore over five seeds and all 15 MVTec AD categories, within 0.003 of anomalib, and a study of what happens to legal search when the EU AI Act is amended under it. Both are written with Murad Eyvazov.
Is my model getting worse?model-nerf-watch scores a model on fixed probes every day with a confidence interval. The honest result so far is how little 60 probes can see: a 20 point drop, not a 5 point one.
Offline AI for accessibility.Live captions that never touch the network, and one short answer from a camera for blind users. Next step for both is real recordings and real photos instead of simulated ones.
Azerbaijani language models. A multilingual hallucination detector on a Raspberry Pi, a benchmark, and the Meristem fine-tune, which I still have to score properly against its base model.
Tools I use
Almost everything here is Python with PyTorch, and C++ when a kernel or a simulation needs it. Tabular work is scikit-learn and LightGBM, language models go through Transformers, PEFT and Docker for deployment, with FastAPI or Streamlit in front and PostgreSQL behind. Every repo runs on GitHub Actions and re-derives its own numbers. On the hardware side I use Raspberry Pi and Arduino, and for local models Ollama and MLX on Apple Silicon.
Story so far
Adscreen
Head of IT
June 2024 to present, Kortrijk, Belgium
Real-time ad targeting pipelines in Python and TensorFlow across a fleet of over 1,000 edge devices. Containerised deployments with Docker and Kubernetes, with monitoring and automated recovery for the computer vision analytics service.
Howest
AI developer and student researcher
September 2024 to June 2025, Kortrijk
Predictive models and transformer NLP pipelines for industry partner projects. Computer vision with CNNs, YOLO and OpenCV, C++ simulation modules, and cross-validation protocols so partner results were reproducible.
STEP IT Academy
Programming instructor
April 2019 to December 2023, Baku
Taught programming and AI to over 200 students and supervised capstone projects from preprocessing to deployment. Built microcontroller prototypes with embedded ML models that became reference coursework.
RED RUSH and ARENA FEST
Founder and team lead
2024, Baku
Founded two events, then moved to a remote advisory role. Led over 30 volunteers at Red Bull F1 and Cars and Coffee events.
Education
Howest University of Applied Sciences
Bachelor, Creative Technologies and AI
Expected 2027, Kortrijk
Advanced deep learning, computer vision, reinforcement learning, MLOps and AI ethics.
STEP IT Academy Azerbaijan
Diploma, programming and IT
2017 to 2023, Baku
International STEM Olympiad
Gold medal, mathematics
2022
Kaggle
IEEE-CIS Fraud Detection, 0.9086 private AUC
Time-ordered folds with a 30-day embargo, so validation never sees the future
I built the autoencoder, DDPM and DDIM myself to see why latent diffusion caught on. Running diffusion in a 4x compressed latent trained 5.3x faster than in pixel space, and at the same number of steps it scored 6.2x better.
PPO with a learned reward model, plus DPO, GRPO, RLOO and Best-of-N, all written from scratch. The interesting part is reward hacking: the proxy reward kept going up to +9.97 while the real objective peaked, then fell to -1.43, worse than where the policy started.
Flow matching and reflow from scratch. One round of reflow made the paths about 15000x straighter, which means you can sample in a single step at 128x less compute without losing quality I could measure.
FlashAttention style attention, timed on my own GPU rather than quoted from the paper. Fusing the kernel is worth about 3x, and it takes me from 22% of peak throughput to roughly 67%.
Optimal transport, Schrodinger bridges trained with IPF, and bridge matching. In my runs bridge matching came out 1.3x to 12x better, and it cost about a quarter of the compute.
DeepSeek-V2's multi-head latent attention, including the decoupled RoPE trick. I checked that the fast inference path matches the naive one (to 6e-07). The payoff is a 14.2x smaller KV cache: 8.4 GB instead of 120 GB at 128k context.
A Dreamer style world model on a pendulum where the agent can't see velocity. With a small budget (57,600 steps) none of my world models beat a random policy. Given the same 384,000 steps as the model-free baseline, all three did, but it cost about 60 times the compute.
Four ways for a robot policy to output an action, on demos that go around an obstacle either left or right. A plain regression head averages the two and drives into the wall for 8.44 of 24 steps. Flow matching gets that down to 3.56 and runs 90x faster than diffusion.
People keep asking if their model got quietly worse, so I built a daily check with confidence intervals. The first thing it showed is how little a small probe set can see: with 60 probes you'll catch a 20 point drop but not a 5 point one. It did catch a planted downgrade to a half-size model (p of 1e-08), and the README says how many probes you need for smaller drops.
Live captions for deaf and hard of hearing people that never go online. On an M4, whisper turbo gets 5.1% word error on clean speech and 11.0% in a noisy room, about 1.7 seconds after you stop talking. The tiny model is much worse with accents and noise, and at the loudest level it starts making up words. There's a test that fails if anything opens a network connection.
For blind users: point the camera, ask one question, and get just the answer, offline on a 4B vision model. A short, strict prompt did better than the usual "describe everything" one (68% against 62%) and stopped the model talking before the answer. Dials and 7-segment displays are still bad, and the test images are synthetic for now, which is the next thing I want to fix.
I planted a shortcut in an anomaly detection dataset to see if the model would cheat. The score went up to 0.998 AUROC, but the heatmaps pointed at the real defect less than half as often (0.75 down to 0.37). It turned out not to be a label shortcut at all. I got the same result on real MVTec images with two detectors and two backbones.
How often does a vision model say an object is there when it isn't? Rewording the question barely changed accuracy, but slipping the object into the premise more than doubled hallucinations, from 6.0% to 13.4%. A CLIP check cut that by 25 to 40%, at some cost in recall.
LoRA fine-tuning Qwen2.5-1.5B to pull fields out of expense receipts, run 15 times on rented GPUs. Exact match went from 48.9% to anywhere between 66.7% and 82.2% depending on the seed, which says a lot about single-run results. One epoch did as well as three.
I tried six reward functions on a 2-link robot arm, and the agent found ways to game two of them. With the distance reward it learned to crash on purpose, every single time. The final PPO agent reached its target less often than a simple PD controller (43.2% against 73.5%) but also crashed less.
Do the code search tools in Google's Gemma 4 agent competition actually point at the code a fix changes? I checked on the training tasks with no model involved. Plain grep found an edited file in its top five for 52 tasks, the built-in embedding search for only 37, and BM25 without test files for 82. Submitted to the competition's paper track.
Searching for a CIFAR-10 network small enough for a microcontroller (50k parameters, 250 KB). The evolved model reached 76.4% against 71.8% for my hand-written baseline, on all 3 seeds, while 33 random designs found nothing better. The C export matches PyTorch.
Checks each claim in a RAG answer against its sources, in any language: pip install groundedness. It uses my own XLM-RoBERTa model, trained on 30 languages, running on a Raspberry Pi 5 in Baku and on Hugging Face, with a paper on SSRN. Across 31 languages it caught 352 of 372 planted errors, where Vectara HHEM caught 265.
My Azerbaijani chat model: Qwen2.5-3B fine-tuned with QLoRA on about 168k instruction pairs in Azerbaijani, Russian, English and Turkish. It's a 1.9 GB file, so you can run it locally with ollama run aghasalim/siba-meristem-1.0. It's on Hugging Face and Ollama.
A benchmark of 2,148 Azerbaijani questions, on Hugging Face. It showed me something about my own model: the fine-tune does better on exam questions than its base (40.0% against 36.9%) but worse at reading comprehension. On one subset, neither beats just answering D every time.
An MCP connector that checks a claim against something real: an actual quote, an actual citation or actual code output, without asking a second language model to judge. It works with Claude and Gemini.
Question answering over the EU AI Act, graded by a different model from the one that answers. On 45 questions I wrote by hand, the right passage is in the top 6 for 90.9% of them, and 90.2% of answers stick to the sources.
Fraud detection where validation respects time, because a random split lets the model peek at the future. That peeking is worth 0.1044 AUC. My private leaderboard score was 0.9086.
Aghasalim Mustafazada and Murad Eyvazov. SSRN, 2026. We wanted to see what happens to a legal search tool when the law it was built on gets amended. So we asked it 11 questions that only the new version of the AI Act can answer. Searching the old 2024 text, it still found the right article most of the time, but it never once found the current wording. Splitting the law by paragraph instead of by article helped (recall went from 69.7% to 75.8%). We also graded 41 answers by hand to check the automatic judge, and it agreed with us on 83% of them. Everything, including all 123 test questions, is in the eu-ai-act-rag repository.
Aghasalim Mustafazada and Murad Eyvazov. SSRN, 2026. I rebuilt PatchCore from scratch and wanted to know if my numbers really hold up. Over all 15 MVTec AD categories and five random seeds, it lands within 0.003 of anomalib, the standard library. The more useful part was a bug in my own threshold. It was meant to give 1% false alarms, but only did in 3 of 15 categories. With proper calibration it's 13 of 15. Code is in the explainable-defect-detector repository.
Aghasalim Mustafazada and Teymur Eyvazov. SSRN, 2026. A small model that reads a RAG answer and highlights the parts its sources don't back up. It works in 31 languages and runs on a Raspberry Pi, and we tested it against LLM judges and the English-only detectors most people use. Code, data and weights are in the groundedness repository.
SSRN preprintRead on SSRN DOI 10.2139/ssrn.7497398
Aghasalim Mustafazada. ARC Prize 2026 paper track, Kaggle, 2026. My program search solved 39 of 1,000 training puzzles and none of the 120 evaluation ones. The write-up is about why. On the evaluation set it never came up with a single candidate, so the verifier had nothing to check. And of the 42 programs it did find, 3 fit every example and were still wrong. Code and numbers are in the arc-prize-2026 repository.
Competition write-up, submittedRead on GitHub DOI 10.5281/zenodo.23003614 (code)