Head to head: muse-spark-1.3 vs Phi-4-mini-instruct
muse-spark-1.3 vs Phi-4-mini-instruct
By Ryan Merket · Published
This matchup tests whether a compact instruction model can keep pace across structured data, coding, localization, editing, math, and constraint-heavy writing. The decisive difference is not style or verbosity, but dependable execution of exact requirements.
muse-spark-1.3 wins this matchup outright: **110.0 to 43.0**, with a **36–0 task record** and no ties. The statistical verdict confirms the result at **limited confidence**, making this a clear win rather than a marginal edge. The separation is most damaging on tasks where correctness and formatting are inseparable. muse-spark returned bare valid JSON, SQL-only answers, properly structured schedules, exact word counts, and the requested `mm:ss` output. Phi-4-mini-instruct repeatedly added forbidden Markdown fences, missed line and word constraints, or produced the right underlying calculation in the wrong format. More seriously, Phi-4-mini-instruct also failed on substance: nonexistent SQL columns, incorrect ranking logic, omitted or mistyped inventory records, broken concurrency fixes, and nonfunctional LFU-cache implementations. Its customer emails, Mexican Spanish localization, and proofreading were likewise less reliable, with awkward phrasing, fabricated edit notes, and occasionally incorrect claims about the supplied text. muse-spark was not flawless at the sentence level—judges noted a few slightly forced or less-than-ideal phrasings—but those were cosmetic blemishes on otherwise correct outputs. **Final call: muse-spark-1.3 is the unambiguous winner, combining stronger reasoning with dramatically better instruction compliance than Phi-4-mini-instruct.**
SQL window query
Given Postgres tables `orders(id, customer_id, created_at, total_cents)` and `customers(id, name)`, write ONE SQL query returning each customer's name and their SECOND-largest order total (in dollars, 2 decimals), excluding customers with fewer than two orders. Order by that value descending. Return ONLY the SQL.
OpenAI: GPT-5.6 Sol Pro: Model A correctly ranks each customer’s orders with ROW_NUMBER, selects the second order, converts cents to dollars with two-decimal rounding, and returns only SQL. Model B references a nonexistent customers.customer_id column, may conflate customers sharing a name, uses distinct-value ranking rather than selecting the second order, and violates the output-format instruction with a Markdown fence. (Second judge pass, order swapped — scores are the average of both: Model A correctly uses ROW_NUMB | Anthropic: Claude Fable 5.1: Model A is a correct, compact query: ROW_N…
Customer delay reply
Write a customer-support email reply to Maya Chen, operations manager at Bristle & Pine Studio. She ordered 240 custom packaging inserts for a product shoot. We promised delivery by Thursday, but a die-cutting machine failure means the shipment will arrive Monday instead. Audience: a frustrated business customer. Tone: accountable, calm, and practical. Length: 140-180 words. Include: a plain-language apology, the new ETA, one concrete goodwill gesture, and a direct phone contact line for urgent coordination. Do not sound legalistic or overly promotional.
OpenAI: GPT-5.6 Sol Pro: Model A clearly explains the delay, accepts responsibility, gives a precise Monday ETA, offers a concrete refund, and provides a direct contact line in a calm, practical email within the requested length. Model B includes awkward and unclear language, describes the delay inaccurately as being pushed “forward,” offers an incoherent shipping concession, and incorrectly signs the sender as an employee of the customer’s company. (Second judge pass, order swapped — scores are the average of both: Mod | Anthropic: Claude Fable 5.1: Model A delivers a clean, accountable emai…
Localize app notice to Mexican Spanish
Translate this app notice into Mexican Spanish for everyday users. Keep the meaning and nuance intact. Tone: friendly and clear, not formal. Length: 2 sentences, maximum 24 words total. "Your draft was saved on this device, but it hasn't been uploaded yet. Connect to the internet to back it up and access it from your other devices."
OpenAI: GPT-5.6 Sol Pro: Model A preserves the meaning in friendly, natural Mexican Spanish and meets the two-sentence, 24-word limit exactly. Model B exceeds the limit, changes “device” to “computer,” and uses less natural wording for uploading and backing up. (Second judge pass, order swapped — scores are the average of both: Model A accurately preserves the meaning in natural, friendly Mexican Spanish and meets the two-sentence, 24-word limit. Model B exceeds the limit, narrows “device” to “computer,” and uses less n | Anthropic: Claude Fable 5.1: Model A delivers a natural, friendly Mexic…
Unit-aware math
A pump moves 3.5 liters every 8 seconds. A tank holds 0.9 cubic meters. Starting empty, how long to fill it, in minutes and seconds (mm:ss), rounded to the nearest second? Show the key steps, then give the final answer on its own line.
OpenAI: GPT-5.6 Sol Pro: Model A is correct, concise, shows the key steps, and provides the requested mm:ss result on its own line. Model B reaches the correct value but does not present the final answer in the explicitly requested mm:ss format and is unnecessarily verbose. (Second judge pass, order swapped — scores are the average of both: Model A is correct, concise, shows the key unit-aware steps, and provides the final answer alone in the requested mm:ss format. Model B reaches the correct rounded time but includes | Anthropic: Claude Fable 5.1: Both compute the correct 34 min 17 s, but M…
Strict JSON extraction
Extract every meeting from this text into a JSON array of {"title","day","start24","durationMinutes"}. Use 24-hour "HH:MM" for start24, integer minutes for duration, and title-case titles. Return ONLY valid JSON, no commentary. Text: "Standup is daily at 9am for a quarter hour. The design review runs thursday from 2:30-4pm. Payroll sync — first monday, 11:00 to noon."
OpenAI: GPT-5.6 Sol Pro: Model A provides valid JSON only and accurately preserves the meeting titles, recurrence labels, times, and durations. Model B adds prohibited Markdown fencing and unnecessarily changes "Standup" to "Daily Standup" and "Daily" to "Every day." (Second judge pass, order swapped — scores are the average of both: Model A preserves the meeting titles, accurately normalizes all times and durations, and returns only valid JSON. Model B adds “Daily” to the Standup title and wraps the array in a Markdown | Anthropic: Claude Fable 5.1: Model A returns bare valid JSON exactly as…
Micro bio with strict rules
Write a speaker bio for Dr. Elian Voss, a hydrologist presenting at a regional water-reuse summit. Follow every rule exactly: - Exactly 4 lines - Line 1: full name only - Line 2: job title and organization - Line 3: exactly 9 words describing expertise - Line 4: one complete sentence, 12-16 words, mentioning Phoenix and 2023 - Entire output must be 38-46 words total - Do not use the words "passion", "leader", or "innovative" - Do not use bullet points, numbering, or quotation marks
OpenAI: GPT-5.6 Sol Pro: Model A satisfies every structural and word-count requirement, although the final word feels slightly forced. Model B uses one line instead of four, misses the required line-specific constraints and total word count, and relies on awkward sentence fragments. (Second judge pass, order swapped — scores are the average of both: Model A satisfies the four-line format, exact line-specific word counts, total 38-word requirement, and all prohibited-content rules, though “today” is slightly awkward. Mod | Anthropic: Claude Fable 5.1: Model A satisfies every constraint: four l…
Proofread a field report
Proofread the passage below for grammar, clarity, punctuation, and consistency WITHOUT changing meaning. Then add 3 brief bullet points noting the main changes. "On Tuesday 14 May, the site team were preparing to replace the west loading-bay sensor, which had been intermittently failing since April. Because the spare arrived late, and the forklift battery wasn't fully charged, the work started nearly an hour behind schedule. Samir checked the wiring diagrams twice, however the printed revision was outdated and didnt match what was installed. We paused, took photos, and called L. Ortega in engineering who confirms the panel had been modified last winter. After that, the replacement itself went smoothly and testing show the alarm cleared. No stock was damaged, but dispatch said the delay have caused two outbound pallets to miss the 16:30 truck."
OpenAI: GPT-5.6 Sol Pro: Model A corrects the comma splice, verb tense and agreement, apostrophe, and punctuation while preserving the report’s meaning and providing three accurate change notes. Model B leaves a comma splice around “however” and inaccurately claims to have corrected the use of “Wednesday,” which does not appear in the passage. (Second judge pass, order swapped — scores are the average of both: Model A correctly fixes the comma splice around “however,” preserves the original date style, and accurately su | Anthropic: Claude Fable 5.1: Model A correctly fixes every error (comma…
Normalize inventory notes to JSON
Convert the messy notes into valid JSON only. Output an array of objects sorted by sku ascending. Schema for each object: {"sku":string,"name":string,"qty":integer,"unit_price":number,"tags":[string],"discontinued":boolean}. Rules: trim whitespace; uppercase sku; qty must be integer; unit_price as number with 2 decimals; split tags on / or comma, lowercase them, remove duplicates, preserve original order; if tags missing use []; interpret discontinued values y/yes/true as true and n/no/false/blank as false. Notes: 1) sku= bk-204 ; name: "Matte Black Folio" ; qty 12 ; price $19.9 ; tags=office / paper, office ; discontinued=no 2) NAME=Glass Clip Set| sku : gc-9 | qty=003 | price=4 | tags=desk/minimal | discontinued=Y 3) sku= rt-77 ; name=River Tile Sample ; qty= 8 ; price= 2.345 ; discontinued= false 4) qty=21; price=$0.99; sku=AL-3; name= "Label Dots" ; tags=warehouse, small-parts, warehouse ; discontinued=
OpenAI: GPT-5.6 Sol Pro: Model A correctly normalizes all four entries, preserves the required fields, and sorts them by SKU. Model B changes BK-204 to BC-204, omits RT-77 entirely, and violates the JSON-only requirement by using a Markdown code fence. (Second judge pass, order swapped — scores are the average of both: Model A correctly includes and normalizes all four records, while Model B omits the RT-77 record, mistypes BK-204 as BC-204, and adds Markdown fences instead of returning JSON only.) | Anthropic: Claude Fable 5.1: Model A produces all four records with correct SKU uppercasing,…
Concurrency bug fix
This TypeScript function is meant to memoize an async loader but has a race: concurrent callers can each trigger the underlying fetch. Fix it so the fetch runs at most once per key, and a rejected fetch does NOT poison the cache (a later call must retry). Return ONLY the corrected function. ```ts const cache = new Map<string, Promise<string>>(); async function load(key: string, fetcher: (k: string) => Promise<string>) { if (cache.has(key)) return cache.get(key)!; const value = await fetcher(key); cache.set(key, Promise.resolve(value)); return value; } ```
OpenAI: GPT-5.6 Sol Pro: Model A correctly caches the in-flight promise and conditionally removes it after rejection, allowing later retries without deleting a newer entry. Model B still awaits before caching, stores a string in a promise-typed map, and violates the requested output format with a code fence. (Second judge pass, order swapped — scores are the average of both: Model A caches the in-flight promise before awaiting it and conditionally removes it after rejection, while Model B both preserves the race and int | Anthropic: Claude Fable 5.1: Model A correctly stores the in-flight pro…
LFU cache with aging
Implement an LFU cache in Python 3. Requirements: class LFUCache with methods __init__(capacity:int), get(key:int)->int, put(key:int,value:int)->None. Evict the key with lowest frequency; if tied, evict the least recently used among that frequency. Both get and put must run in O(1) average time. Capacity may be 0. Frequencies start at 1 on insert and increase on successful get or updating an existing key with put. Use only the standard library. Return code only, no explanation, no tests.
OpenAI: GPT-5.6 Sol Pro: Model A correctly implements O(1)-average LFU/LRU eviction, frequency updates, and zero-capacity handling. Model B is nonfunctional due to undefined instance attributes, invalid sentinel/list usage, broken node manipulation and eviction, and Markdown fences despite the code-only instruction. (Second judge pass, order swapped — scores are the average of both: Model A correctly implements O(1)-average LFU/LRU eviction and handles updates and zero capacity, while Model B uses undefined attributes, | Anthropic: Claude Fable 5.1: Model A delivers a correct, idiomatic O(1)…
Precise proofreading
Correct ONLY the grammar, spelling, and punctuation errors in the sentence below without changing wording, tone, or meaning, then on a second line list each change as 'was -> now'. Sentence: "Its been a long time since we seen there new office, and me and him was suprised how quite it was."
OpenAI: GPT-5.6 Sol Pro: Model A corrects every grammatical, spelling, and punctuation error while preserving the meaning and lists each change in the requested format. Model B leaves most errors in the sentence, includes a nonexistent change, and lists claimed corrections that were not applied. (Second judge pass, order swapped — scores are the average of both: Model A correctly fixes every error and accurately lists each change on the second line. Model B leaves most errors in the sentence, includes a meaningless “see | Anthropic: Claude Fable 5.1: Model A correctly fixes every error (It's,…
Constraint scheduling
Four talks (A, B, C, D) fill four 1-hour slots 9,10,11,12. Constraints: A is before D; C is not first; B is immediately after A; D is not at 12. Give the ONE valid schedule as 'slot: talk' lines, then a one-line justification. If impossible, say so and explain.
OpenAI: GPT-5.6 Sol Pro: Model A gives the correct unique schedule in the requested line-by-line format and concisely explains why it is forced. Model B has the correct ordering, but it does not use the required “slot: talk” lines, includes a stray heading, and does not justify uniqueness. (Second judge pass, order swapped — scores are the average of both: Model A gives the unique correct schedule in the requested line-by-line format and provides a concise justification of uniqueness. Model B reaches the same correct sc | Anthropic: Claude Fable 5.1: Both arrive at the unique correct schedule…
Matchup powered by OpenRouter.