<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Miikael Mändmets</title><link>https://miikaelm.github.io/</link><description>Recent content on Miikael Mändmets</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Thu, 30 Jul 2026 10:00:00 +0300</lastBuildDate><atom:link href="https://miikaelm.github.io/index.xml" rel="self" type="application/rss+xml"/><item><title>Nobody checked whether the judge can see: auditing VLM-as-judge for text style editing</title><link>https://miikaelm.github.io/posts/nobody-checked-whether-the-judge-can-see/</link><pubDate>Thu, 30 Jul 2026 10:00:00 +0300</pubDate><guid>https://miikaelm.github.io/posts/nobody-checked-whether-the-judge-can-see/</guid><description>&lt;p&gt;&lt;strong&gt;Main takeaways&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;What I found:&lt;/em&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Judge reliability is judge-dependent, not just task-dependent: on identical stimuli, Gemini 3.1 Pro&amp;rsquo;s scores track error magnitude on 11 of 12 dimensions, GPT-4o&amp;rsquo;s on 6, and Qwen2.5-VL&amp;rsquo;s on 3.&lt;/li&gt;
&lt;li&gt;Where judges do penalise errors, they start late: onsets sit at or beyond the magnitude that&amp;rsquo;s clearly visible to a human, and several error types are never reliably penalised at all.&lt;/li&gt;
&lt;li&gt;Verifying an instructed value (is this the right hex colour, the right angle?) is a harder task for a judge than noticing an incidental anomaly. Colour verification fails outright for every judge tested, at every magnitude.&lt;/li&gt;
&lt;li&gt;Holistic &amp;ldquo;overall quality&amp;rdquo; scores diverge sharply across judges on identical, unedited images (a 2.5-point spread on a 5-point scale) while decomposed rubric dimensions do not (0.41 points). Controlled-stimulus evidence that decomposed rubrics are the right design choice.&lt;/li&gt;
&lt;li&gt;Judges can see errors they don&amp;rsquo;t penalise: shown the same images under a reference-based framing instead of the benchmark&amp;rsquo;s instruction-based one, GPT-4o registers significantly more errors than it penalises.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Contents&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>What building an evaluation taught me about reading them</title><link>https://miikaelm.github.io/posts/what-building-an-evaluation-taught-me/</link><pubDate>Thu, 30 Jul 2026 09:00:00 +0300</pubDate><guid>https://miikaelm.github.io/posts/what-building-an-evaluation-taught-me/</guid><description>&lt;p&gt;I recently published &lt;a href="https://miikaelm.github.io/posts/nobody-checked-whether-the-judge-can-see/"&gt;an audit&lt;/a&gt; of whether VLM judges can perceive the errors that text-in-image editing benchmarks ask them to penalise. This was my first scientific work, written under a thesis deadline, with not much prior work to build on. Most design decisions were mine to make and I had little to check them against. Below is what the process taught me.&lt;/p&gt;
&lt;h2 id="the-thesis-did-not-start-as-an-audit"&gt;The thesis did not start as an audit&lt;/h2&gt;
&lt;p&gt;The plan was a synthetic training pipeline for quantitative text style editing. Render designed documents from HTML/CSS in a headless browser, apply edits programmatically (&amp;ldquo;scale the title by 20%&amp;rdquo;, &amp;ldquo;shift the byline 40 pixels right&amp;rdquo;), and train on the resulting code-to-image pairs, where the correct answer is known to the pixel. The capability gap was somewhat documented: &lt;a href="https://arxiv.org/abs/2512.16270"&gt;TextEditBench&lt;/a&gt; reports that models handle content edits well but fail on spatial and stylistic transformations. Rendered supervision looked like an obvious way at it that nobody had taken.&lt;/p&gt;</description></item><item><title>About</title><link>https://miikaelm.github.io/about/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://miikaelm.github.io/about/</guid><description>About</description></item></channel></rss>