Hi, I'm Aghasalim Mustafazada

Aghasalim Mustafazada

I'm a third year AI student at Howest in Kortrijk, Belgium, originally from Baku. I spend most of my time on anomaly detection and model evaluation, and I don't trust a number until I have reproduced it myself. Most of what is on this page is either a known system rebuilt from the paper and measured on my own laptop, or an evaluation where the obvious story turned out to be wrong.

The numbers below all come from runs in the linked repos, and most of those repos fail their own CI if the README stops matching the results. Outside the from-scratch work I built a multilingual hallucination detector that runs on a Raspberry Pi, a benchmark for Azerbaijani LLMs, and offline accessibility tools for deaf and blind users. Before Belgium I taught programming and AI to 200+ students in Baku, and I run the IT side of a digital signage company with 1,000+ edge devices.

Beyond work, I like mathematics olympiads (a gold medal in 2022), events (I founded two in Baku) and cycling around Kortrijk.

Want to work together, or just say hi?

GitHub Kaggle LinkedIn SSRN paper Hugging Face Scholar

NovDecJanFebMarAprMayJunJulAugSepOct
3,480 contributions in the last year, across my two GitHub accounts, aghasalim and MustafazadaAghasalim.Less More

Right now

Last updated 10 October 2026.

Tools I use

Almost everything here is Python with PyTorch, and C++ when a kernel or a simulation needs it. Tabular work is scikit-learn and LightGBM, language models go through Transformers, PEFT and Docker for deployment, with FastAPI or Streamlit in front and PostgreSQL behind. Every repo runs on GitHub Actions and re-derives its own numbers. On the hardware side I use Raspberry Pi and Arduino, and for local models Ollama and MLX on Apple Silicon.

Story so far

Education

Curious? Check out my projects.

latent-diffusion-from-scratch

I built the autoencoder, DDPM and DDIM myself to see why latent diffusion caught on. Running diffusion in a 4x compressed latent trained 5.3x faster than in pixel space, and at the same number of steps it scored 6.2x better.

autoencoderddimddpmdiffusion-modelsgenerative-models

rlhf-ppo-from-scratch

PPO with a learned reward model, plus DPO, GRPO, RLOO and Best-of-N, all written from scratch. The interesting part is reward hacking: the proxy reward kept going up to +9.97 while the real objective peaked, then fell to -1.43, worse than where the policy started.

alignmentdpogoodharts-lawgrpollm

rectified-flow-from-scratch

Flow matching and reflow from scratch. One round of reflow made the paths about 15000x straighter, which means you can sample in a single step at 128x less compute without losing quality I could measure.

diffusion-modelsflow-matchinggenerative-modelsode-solverspytorch

flash-attention-from-scratch

FlashAttention style attention, timed on my own GPU rather than quoted from the paper. Fusing the kernel is worth about 3x, and it takes me from 22% of peak throughput to roughly 67%.

attentioncudaflash-attentiongpu-kernelsperformance-optimization

schrodinger-bridge-from-scratch

Optimal transport, Schrodinger bridges trained with IPF, and bridge matching. In my runs bridge matching came out 1.3x to 12x better, and it cost about a quarter of the compute.

bridge-matchingdiffusion-modelsentropic-optimal-transportgenerative-modelsoptimal-transport

mla-from-scratch

DeepSeek-V2's multi-head latent attention, including the decoupled RoPE trick. I checked that the fast inference path matches the naive one (to 6e-07). The payoff is a 14.2x smaller KV cache: 8.4 GB instead of 120 GB at 128k context.

attentiondeepseekinference-optimizationkv-cachellm

world-model-from-scratch

A Dreamer style world model on a pendulum where the agent can't see velocity. With a small budget (57,600 steps) none of my world models beat a random policy. Given the same 384,000 steps as the model-free baseline, all three did, but it cost about 60 times the compute.

dreamerimaginationmodel-based-rlpomdppytorch

vla-from-scratch

Four ways for a robot policy to output an action, on demos that go around an obstacle either left or right. A plain regression head averages the two and drives into the wall for 8.44 of 24 steps. Flow matching gets that down to 3.56 and runs 90x faster than diffusion.

action-chunkingdiffusion-policyflow-matchingimitation-learningpytorch

model-nerf-watch

People keep asking if their model got quietly worse, so I built a daily check with confidence intervals. The first thing it showed is how little a small probe set can see: with 60 probes you'll catch a 20 point drop but not a 5 point one. It did catch a planted downgrade to a half-size model (p of 1e-08), and the README says how many probes you need for smaller drops.

benchmarkingllm-evaluationmodel-driftollamastatistics

offline-live-captions

Live captions for deaf and hard of hearing people that never go online. On an M4, whisper turbo gets 5.1% word error on clean speech and 11.0% in a noisy room, about 1.7 seconds after you stop talking. The tiny model is much worse with accents and noise, and at the loudest level it starts making up words. There's a test that fails if anything opens a network connection.

accessibilitydeafhard-of-hearingmlxon-device

one-answer-vision

For blind users: point the camera, ask one question, and get just the answer, offline on a 4B vision model. A short, strict prompt did better than the usual "describe everything" one (68% against 62%) and stopped the model talking before the answer. Dials and 7-segment displays are still bad, and the test images are synthetic for now, which is the next thing I want to fix.

accessibilityblindcomputer-visionlow-visionollama

spurious-ad

I planted a shortcut in an anomaly detection dataset to see if the model would cheat. The score went up to 0.998 AUROC, but the heatmaps pointed at the real defect less than half as often (0.75 down to 0.37). It turned out not to be a label shortcut at all. I got the same result on real MVTec images with two detectors and two backbones.

anomaly-detectioncomputer-visionexplainable-aimodel-evaluationpytorch

vlm-hallucination-eval

How often does a vision model say an object is there when it isn't? Rewording the question barely changed accuracy, but slipping the object into the premise more than doubled hallucinations, from 6.0% to 13.4%. A CLIP check cut that by 25 to 40%, at some cost in recall.

blipcliphallucination-detectionmodel-evaluationmultimodal

lora-forgetting

LoRA fine-tuning Qwen2.5-1.5B to pull fields out of expense receipts, run 15 times on rented GPUs. Exact match went from 48.9% to anywhere between 66.7% and 82.2% depending on the seed, which says a lot about single-run results. One epoch did as well as three.

catastrophic-forgettingfine-tuninginformation-extractionllmlora

rl-arm-reward-shaping

I tried six reward functions on a 2-link robot arm, and the agent found ways to game two of them. With the distance reward it learned to crash on purpose, every single time. The final PPO agent reached its target less often than a simple PD controller (43.2% against 73.5%) but also crashed less.

gymnasiumpporeinforcement-learningreward-hackingreward-shaping

gemma4-code-graph-localization

Do the code search tools in Google's Gemma 4 agent competition actually point at the code a fix changes? I checked on the training tasks with no model involved. Plain grep found an edited file in its top five for 52 tasks, the built-in embedding search for only 37, and BM25 without test files for 82. Submitted to the competition's paper track.

gemma-4code-graphbug-localization

enas-microcontroller

Searching for a CIFAR-10 network small enough for a microcontroller (50k parameters, 250 KB). The evolved model reached 76.4% against 71.8% for my hand-written baseline, on all 3 seeds, while 33 random designs found nothing better. The C export matches PyTorch.

neural-architecture-searchmicrocontrollercifar-10

groundedness

Checks each claim in a RAG answer against its sources, in any language: pip install groundedness. It uses my own XLM-RoBERTa model, trained on 30 languages, running on a Raspberry Pi 5 in Baku and on Hugging Face, with a paper on SSRN. Across 31 languages it caught 352 of 372 planted errors, where Vectara HHEM caught 265.

hallucination-detectionragmultilingualxlm-roberta

siba-meristem-1.0

My Azerbaijani chat model: Qwen2.5-3B fine-tuned with QLoRA on about 168k instruction pairs in Azerbaijani, Russian, English and Turkish. It's a 1.9 GB file, so you can run it locally with ollama run aghasalim/siba-meristem-1.0. It's on Hugging Face and Ollama.

qwen2.5qloraazerbaijanigguf

azbench

A benchmark of 2,148 Azerbaijani questions, on Hugging Face. It showed me something about my own model: the fine-tune does better on exam questions than its base (40.0% against 36.9%) but worse at reading comprehension. On one subset, neither beats just answering D every time.

azerbaijanillm-benchmarkbelebeleinclude

groundcheck-mcp

An MCP connector that checks a claim against something real: an actual quote, an actual citation or actual code output, without asking a second language model to judge. It works with Claude and Gemini.

ai-safetyclaudefact-checkinggeminihallucination-detection

eu-ai-act-rag

Question answering over the EU AI Act, graded by a different model from the one that answers. On 45 questions I wrote by hand, the right passage is in the top 6 for 90.9% of them, and 90.2% of answers stick to the sources.

eu-ai-acthybrid-searchinformation-retrievalllm-evaluationnlp

ieee-fraud-ml

Fraud detection where validation respects time, because a random split lets the model peek at the future. That peeking is worth 0.1044 AUC. My private leaderboard score was 0.9086.

feature-engineeringfraud-detectiongradient-boostingkaggletabular-data

Publications

Retrieval over an Amended Law: Corpus Versioning, Paragraph-Level Citation and Judge Validation for Question Answering on the EU AI Act

Aghasalim Mustafazada and Murad Eyvazov. SSRN, 2026. We wanted to see what happens to a legal search tool when the law it was built on gets amended. So we asked it 11 questions that only the new version of the AI Act can answer. Searching the old 2024 text, it still found the right article most of the time, but it never once found the current wording. Splitting the law by paragraph instead of by article helped (recall went from 69.7% to 75.8%). We also graded 41 answers by hand to check the automatic judge, and it agreed with us on 83% of them. Everything, including all 123 test questions, is in the eu-ai-act-rag repository.

SSRN preprint, in review SSRN page Read the PDF

Evaluating a From-Scratch PatchCore on MVTec AD: Metric Protocol, Seed Variance, False-Alarm Intervals and Threshold Calibration

Aghasalim Mustafazada and Murad Eyvazov. SSRN, 2026. I rebuilt PatchCore from scratch and wanted to know if my numbers really hold up. Over all 15 MVTec AD categories and five random seeds, it lands within 0.003 of anomalib, the standard library. The more useful part was a bug in my own threshold. It was meant to give 1% false alarms, but only did in 3 of 15 categories. With proper calibration it's 13 of 15. Code is in the explainable-defect-detector repository.

SSRN preprint, in review SSRN page Read the PDF

A Multilingual Groundedness Detector that Runs on a Raspberry Pi: Trained Token Classification against LLM Judges and English-only Detectors in 31 Languages

Aghasalim Mustafazada and Teymur Eyvazov. SSRN, 2026. A small model that reads a RAG answer and highlights the parts its sources don't back up. It works in 31 languages and runs on a Raspberry Pi, and we tested it against LLM judges and the English-only detectors most people use. Code, data and weights are in the groundedness repository.

SSRN preprint Read on SSRN DOI 10.2139/ssrn.7497398

Nothing to Verify: Why a Verified DSL Search Scores Zero on ARC-AGI-2

Aghasalim Mustafazada. ARC Prize 2026 paper track, Kaggle, 2026. My program search solved 39 of 1,000 training puzzles and none of the 120 evaluation ones. The write-up is about why. On the evaluation set it never came up with a single candidate, so the verifier had nothing to check. And of the 42 programs it did find, 3 fit every example and were still wrong. Code and numbers are in the arc-prize-2026 repository.

Competition write-up, submitted Read on GitHub DOI 10.5281/zenodo.23003614 (code)

Contact

Open to AI and backend internships, in Belgium or remote. Email salim.mustafazada@student.howest.be, or message me on LinkedIn.