Head to head: Phi-4-multimodal-instruct vs Codestral-2501
Phi-4-multimodal-instruct vs Codestral-2501
By Ryan Merket · Published
This matchup pits Phi-4-multimodal-instruct’s stronger reasoning and extraction against Codestral-2501’s generally sharper formatting, localization, and code-adjacent work. The task-level results also expose substantial instruction-following flaws on both sides.
Codestral-2501 posts the higher aggregate score, 67.0 to 60.3, but that gap is not decisive: the evaluation assigns only a 50% chance that either model is genuinely better. In practical terms, this is a sample tie—not evidence of a meaningful overall lead. Phi-4-multimodal-instruct was more dependable on precise proofreading, constraint scheduling, field-detail extraction, and the unsatisfiable client-demo puzzle. It tended to preserve consequential details that Codestral dropped, but it was hardly clean: missed corrections, excess explanation, and broken output constraints repeatedly weakened otherwise better answers. Codestral-2501 had the edge on outreach formatting, European Spanish localization, SQL, support-thread summaries, and the weaker TypeScript LRU attempts. Yet many of those wins were merely less-wrong performances: its SQL still selected the wrong rank in places, its localization exceeded the character limit, and its cache implementation sacrificed the required O(1) eviction. The properly implemented Python LRU tasks were ties, with both models needlessly adding Markdown fences. The split results on faithful rewriting and vendor-quote normalization capture the broader problem: small prompt variations changed which model made the critical mistake. **Final call: too close to call. Treat Phi-4-multimodal-instruct and Codestral-2501 as effectively even here, with task fit—not the seven-point aggregate gap—driving the choice.**
Tightly constrained outreach note
Write a customer-success outreach email that follows EVERY rule exactly: - Output exactly 5 lines, in this order: Subject line, greeting, body sentence 1, body sentence 2, sign-off line. - The subject line must start with "Subject:" and contain exactly 6 words after it. - Greeting must be exactly: "Hi Jordan," - Body sentence 1 must be 12-16 words and mention the dashboard migration on March 18. - Body sentence 2 must be 12-16 words and include both a question mark and the exact phrase "reply by Friday". - Do NOT use the words "please", "sorry", or "thanks" anywhere. - Sign-off line must be exactly: "—Nia, Customer Success"
OpenAI: GPT-5.6 Sol Pro: Model B is better because it preserves the required five-line structure and reads naturally, although its subject has only three words, both body sentences are too short, and it omits “reply by Friday.” Model A includes the required phrase, but it adds blank lines, combines both body sentences on one line, uses seven subject words, has undersized body sentences, and is less polished. (Second judge pass, order swapped — scores are the average of both: Model B has the required five-line structure, | Anthropic: Claude Fable 5.1: Both outputs break several hard constraint…
LRU cache in Python
Implement an LRU cache in Python. Return CODE ONLY, no explanation. Requirements: define class LRUCache with methods __init__(capacity: int), get(key: int) -> int, put(key: int, value: int) -> None. Both operations must run in O(1) average time using a hash map plus a doubly linked list; do not use OrderedDict or functools. get returns -1 if missing. When capacity is exceeded, evict the least recently used item. The code must be complete and runnable as-is.
OpenAI: GPT-5.6 Sol Pro: Model A and Model B both correctly implement O(1)-average LRU behavior with a hash map and doubly linked list. Model A and Model B equally violate the code-only requirement by including Markdown fences, but are otherwise complete and equivalent in quality. (Second judge pass, order swapped — scores are the average of both: Model B and Model A both correctly implement O(1)-average LRU caches using a hash map and doubly linked list, including proper recency updates and eviction. Both slightly miss | Anthropic: Claude Fable 5.1: Model A and Model B both deliver correct,…
Extract field visit details
Pull structured facts from this messy field note. Return ONLY JSON with this schema: {"site_id":"","visit_date":"YYYY-MM-DD","start_time":"HH:MM","end_time":"HH:MM","inspector":"","weather":"","issues":[...],"meter_reading_kwh":number,"follow_up_required":true|false}. Use 24-hour time. Include all distinct issues mentioned, normalized to short phrases. Text: "Ops notebook — South basin check. Site: SB-17 / old sticker says S817, ignore that. Tuesday 14 May 2024, arrived 8:07am, wrapped up 10:41 a.m. Inspector: Lena Ortiz (with trainee M. Ko). Weather mostly overcast, light drizzle by the end. Main notes: fence on east side bent inward near panel C; algae buildup around outlet grate; sample cooler latch cracked but still closes. Meter photo shows 18439 kWh — not the smudged 18489 guess from first glance. No spill observed. Need a contractor quote for the fence this week; the cooler can wait until next month."
OpenAI: GPT-5.6 Sol Pro: Model A better preserves the issue-specific locations (“near panel C” and “around outlet grate”) while keeping the phrases concise and normalized. Model B is otherwise accurate but loses useful details from the field note. (Second judge pass, order swapped — scores are the average of both: Model A is better because it preserves the relevant locations of the fence damage and algae buildup while keeping the issues concise. Model B is otherwise correct but omits those useful details.) | Anthropic: Claude Fable 5.1: Both outputs extract every core field correctly (site_id…
Precise proofreading
Correct ONLY the grammar, spelling, and punctuation errors in the sentence below without changing wording, tone, or meaning, then on a second line list each change as 'was -> now'. Sentence: "Its been a long time since we seen there new office, and me and him was suprised how quite it was."
OpenAI: GPT-5.6 Sol Pro: Model A correctly fixes the apostrophe in “It's” and the spelling of “surprised,” but leaves several errors and provides an inaccurate change list; Model B makes no corrections and supplies a nonsensical list. (Second judge pass, order swapped — scores are the average of both: Model B leaves every error uncorrected and provides an incoherent change list. Model A correctly applies the apostrophe in “It’s” and identifies some additional corrections, but fails to apply most of them, omits several e | Anthropic: Claude Fable 5.1: Model A applies only one fix (It's) to the…
LRU cache
Implement a class `LRUCache<K, V>` in TypeScript with a fixed capacity set in the constructor, and O(1) `get(key): V | undefined` and `set(key, value): void`. Accessing or updating a key must mark it most-recently-used; inserting beyond capacity must evict the least-recently-used entry. Return ONLY the class, no prose.
OpenAI: GPT-5.6 Sol Pro: Model B generally implements LRU semantics for positive capacities, but its eviction scan is O(n), violating the required O(1) set operation. Model A also uses O(n) array operations and has serious correctness bugs with falsy values, duplicate recency entries on updates, and falsy keys during eviction. (Second judge pass, order swapped — scores are the average of both: Model B generally preserves LRU semantics but violates the required O(1) eviction by scanning all usage entries and mishandles z | Anthropic: Claude Fable 5.1: Model A has real correctness bugs (falsy-v…
Constraint scheduling
Four talks (A, B, C, D) fill four 1-hour slots 9,10,11,12. Constraints: A is before D; C is not first; B is immediately after A; D is not at 12. Give the ONE valid schedule as 'slot: talk' lines, then a one-line justification. If impossible, say so and explain.
OpenAI: GPT-5.6 Sol Pro: Model A gives the unique valid schedule and accurately verifies every constraint, though the justification is longer than the requested single line. Model B proposes a schedule that violates both “A is before D” and “C is not first,” while incorrectly claiming that it satisfies them. (Second judge pass, order swapped — scores are the average of both: Model B’s proposed schedule violates both “C is not first” and “B is before D,” and its reasoning contains multiple contradictions. Model A gives t | Anthropic: Claude Fable 5.1: Model A gives the unique correct schedule…
Schedule four client demos
Schedule four demo calls — Atlas, Birch, Cinder, and Delta — into four one-hour slots: 09:00, 10:00, 11:00, 12:00. Assign one unique presenter to each call from Imani, Pavel, Sora, and Tarek. Use each slot and presenter exactly once. Return ONLY a table with columns Client | Time | Presenter. Constraints: 1) Atlas is earlier than Birch. 2) Cinder is at 11:00. 3) Tarek presents the 12:00 call. 4) Delta is not at 09:00. 5) Pavel presents Birch. 6) Imani is neither presenting Atlas nor the 10:00 call. 7) Sora presents the client scheduled immediately before Delta. 8) Atlas is not at 10:00. 9) The 09:00 call is not presented by Pavel. 10) Delta is presented by Tarek or Imani.
OpenAI: GPT-5.6 Sol Pro: Model A is closer, satisfying eight of the ten constraints, while Model B violates five constraints. No fully valid schedule exists because the forced timing and presenter assignments ultimately require Imani to present Atlas, contradicting constraint 6; Model A also violates the table-only requirement with its introductory sentence. (Second judge pass, order swapped — scores are the average of both: Model A is substantially closer, violating only constraints 3 and 7, whereas Model B violates co | Anthropic: Claude Fable 5.1: The constraint set is actually unsatisfiab…
Faithful rewrite
Rewrite this blunt internal note as a warm, professional Slack message to the whole engineering team (60–90 words), keeping every fact intact and adding no new commitments. Return only the message. Note: "Deploy is frozen until the memory leak in the image service is fixed. Priya is on it. Do not merge to main. ETA tomorrow noon."
OpenAI: GPT-5.6 Sol Pro: Model A more faithfully keeps the merge restriction tied to resolving the leak, whereas Model B changes the noon ETA into an update deadline and implies merging may resume then. Both are warm and professional but fall below the required 60–90 words, and Model A’s “back on track” wording slightly overstates the ETA. (Second judge pass, order swapped — scores are the average of both: Model A preserves the noon ETA more faithfully, while Model B changes it into a promise of an update by noon. Model | Anthropic: Claude Fable 5.1: Both outputs fall short of the 60–90 word…
Localization with tone
Translate this app onboarding line into natural, friendly European Spanish suitable for a mobile toast (keep it under 60 characters, no exclamation marks): "You're all set — your first backup starts tonight." Return only the translation, then the character count in parentheses.
OpenAI: GPT-5.6 Sol Pro: Model B is accurate, natural European Spanish, but it is 63 characters rather than the stated 69 and exceeds the 60-character limit. Model A is unnatural and inaccurate, exceeds the limit, includes prohibited exclamation marks, and also gives an incorrect character count. (Second judge pass, order swapped — scores are the average of both: Model B is natural and semantically accurate, but its translation is 63 characters rather than the stated 69 and exceeds the 60-character limit. Model A is unn | Anthropic: Claude Fable 5.1: Model A violates the no-exclamation rule,…
SQL window query
Given Postgres tables `orders(id, customer_id, created_at, total_cents)` and `customers(id, name)`, write ONE SQL query returning each customer's name and their SECOND-largest order total (in dollars, 2 decimals), excluding customers with fewer than two orders. Order by that value descending. Return ONLY the SQL.
OpenAI: GPT-5.6 Sol Pro: Model B is executable and correctly excludes customers with fewer than two orders, but it returns the largest total rather than the second-largest. Model A is not executable because `OrderTotals` does not expose the referenced `customer_id`, and its ranking logic also cannot derive the requested result. (Second judge pass, order swapped — scores are the average of both: Model B is valid SQL and correctly excludes customers with fewer than two orders, but it returns the largest total rather than | Anthropic: Claude Fable 5.1: Model A is not even executable (it referenc…
Summarize support thread in bullets
Summarize the support thread below into EXACTLY 4 bullet points. Each bullet must be 11-16 words, start with a bold label followed by a colon, and stay faithful to the thread. Do not mention anything not stated. Thread: Rina (customer), 09:12 — Since yesterday's mobile app update, our drivers can't upload delivery photos on Android tablets. Spinner runs forever after tapping Submit. Devon (support), 09:24 — Thanks. Which app version and Android version? Rina, 09:31 — App 5.18.2. Tablets are Android 11 on Zebra units. iPhones seem fine. Marta (engineering), 10:02 — I can reproduce on Android 11 only when photo filenames contain spaces. Regression likely from the new image sanitizer. Devon, 10:15 — Temporary workaround: rename photo files without spaces before upload. Marta, 11:06 — Fix merged to hotfix branch. QA says 5.18.3 resolves it on Android 11 and 13. Devon, 11:22 — We'll publish 5.18.3 by 3pm UTC and notify affected customers.
OpenAI: GPT-5.6 Sol Pro: Model B is more faithful and clear, though Model B omits bullet markers and its second line exceeds 16 words. Model A also omits bullet markers, lacks required colons, has an overlong first line, and incorrectly says “roman spaces.” (Second judge pass, order swapped — scores are the average of both: Model B is substantially more faithful and uses bold labels with colons, though it lacks bullet markers and one line exceeds 16 words. Model A also lacks bullet markers, omits required colons, exceed | Anthropic: Claude Fable 5.1: Model B stays faithful to the thread, uses…
Normalize vendor quotes to JSON
Convert the messy quote list below into VALID JSON only. Output an object with one key, "quotes", whose value is an array of objects sorted by sku ascending. Each object must have exactly these fields: sku (string), vendor (string), qty (integer), unit_price_usd (number with 2 decimals), lead_time_days (integer), in_stock (boolean). If lead time is given in weeks, convert using 7 days per week. Treat "available now"/"ships today" as 0 days. Messy data: - Vendor Northlight: SKU QX-44, quantity 12, $7.5 each, lead 2 wks, stock yes. - west pier supply :: qx-07 / qty=3 / unit USD 19.00 / available now / in stock true - Northlight -> item QX-07 ; 3 units ; $18.40 ea ; lead 5 days ; stock: no - Alder Trade | sku QX-44 | qty 12 | unit price 7.45 USD | ships today | stock=Y - Alder Trade | sku QX-91 | qty 1 | unit price 104 USD | 1 week lead | stock=N
OpenAI: GPT-5.6 Sol Pro: Model A correctly normalizes the vendor as "Northlight," whereas Model B incorrectly retains the label as part of "Vendor Northlight." Model A writes one price as 7.5 rather than 7.50, and Model A and Model B each use Markdown fences despite the valid-JSON-only requirement. (Second judge pass, order swapped — scores are the average of both: Model A correctly normalizes the vendor name to "Northlight," whereas Model B incorrectly retains the label as "Vendor Northlight." Model A's main flaw is re | Anthropic: Claude Fable 5.1: Model A extracts all data correctly and so…
Matchup powered by OpenRouter.