Head to head: Google: Gemini 3.6 Flash vs gpt-oss-120b

Google: Gemini 3.6 Flash vs gpt-oss-120b

This was competitive, but Gemini 3.6 Flash put more points on the board by being the steadier instruction-follower across practical tasks. gpt-oss-120b had some sharp wins on tightly constrained writing and unit-aware reasoning, yet Gemini’s broader reliability carried the matchup.

The topline is straightforward: **Gemini 3.6 Flash wins on aggregate, 108.8 to 102.8**, with a **75% confidence lean**. That is not a blowout, and it should not be framed like one. But it is a real edge, backed by the task count too: **5 wins for Gemini, 3 for gpt-oss-120b, with 4 ties**. What decided it was not raw brilliance so much as fewer avoidable mistakes. Gemini was better on the kind of work that punishes sloppiness: the customer outage email was more audience-appropriate and less ops-jargony; the scheduling answer matched the requested output format more closely; the phone normalizer handled extension patterns more robustly; and in both LRU cache tasks, Gemini was simply the safer implementation. The Python cache result is especially damaging for gpt-oss-120b: invalid syntax in a code-generation task is the sort of unforced error that swings matchups. To gpt-oss-120b’s credit, its wins were legitimate and specific. It was better at **precise proofreading**, where formatting discipline mattered, and it took the **constrained meeting summary** by actually respecting the tight word-count limits. It also edged the **unit-aware math** task with cleaner, more explicit dimensional reasoning. When the brief is narrow and compliance-heavy, gpt-oss-120b can look very sharp. The tie column matters here too. Several tasks were effectively dead even, including shipment extraction and the SQL window query, and the nuanced classification and incident-note proofreading were split enough across judge passes that neither model established a durable advantage. That keeps this from being a statement win. But across the full set, Gemini was the more dependable model where real users tend to notice failure: runnable code, exact formatting, and tone calibrated to the audience. **Final call: Gemini 3.6 Flash takes this one on consistency, not dominance. It’s a lean win, but a deserved one.**

Nuanced classification

Classify each review's sentiment as "positive", "negative", or "mixed", and give a 6-word-max reason. Return ONLY a JSON array of {"text","label","reason"} in input order. Reviews: ["Fast shipping but the fabric feels cheap.", "Absolutely love it, wearing it daily!", "It broke after a week. Refund was quick and painless though."]

Google: Gemini 3.6 Flash: [ { "text": "Fast shipping but the fabric feels cheap.", "label": "mixed", "reason": "Fast shipping but cheap fabric." }, { "text": "Absolutely love it, wearing it daily!", "label": "positive", "reason": "Loves product and wears it daily." }, { "text": "It broke after a week. Refund was quick and painless though.", "label": "mixed", "reason": "Item broke, but refund was easy." } ]
gpt-oss-120b: [ {"text":"Fast shipping but the fabric feels cheap.","label":"mixed","reason":"Quick delivery, low-quality material."}, {"text":"Absolutely love it, wearing it daily!","label":"positive","reason":"Adore it, uses it every day."}, {"text":"It broke after a week. Refund was quick and painless though.","label":"mixed","reason":"Product failed, refund handled smoothly."} ]

Model B is slightly better because all labels are correct and its reasons are concise, natural, and consistently within the six-word limit. Model A is also strong, but one reason is a bit less polished and more repetitive in phrasing. (Second judge pass, order swapped — scores are the average of both: Model A is slightly better because all reasons are concise, natural, and clearly within the six-word limit while preserving the sentiment cues. Model B is also strong, but one reason is a bit less direct and slightly less polished than Model A’s phrasing.)

Customer outage update email

Write an email to enterprise customers about a resolved outage in a payroll SaaS product. Audience: HR administrators. Tone: calm, accountable, non-technical. Length: 140-180 words. Include: what happened (employees in Canada could not download T4 slips for about 95 minutes this morning), what customers should do now (retry; no action if successful), reassurance about data integrity, and a sincere apology. Do not mention internal systems or root-cause speculation.

Google: Gemini 3.6 Flash: **Subject:** Service Restored: T4 Slip Downloads Now Available Dear HR Administrator, We are writing to inform you that an issue affecting T4 tax slip downloads for Canadian employees earlier today has been fully resolved. For approximately 95 minutes this morning, some employees were unable to access or download their T4 slips through the payroll portal. Our team addressed the issue promptly, and full service has been restored. We want to reassure you that all payroll data remained entirely secure, accurate, and uncompromised throughout this period. **What you should do now:** Please advise your Canadian employees to log in and attempt to download their T4 slips again. If the download is successful, no further action is needed. Should any employee continue to experience difficulties, please have them try again after refreshing their browser, or reach out to our customer support team for assistance. We sincerely apologize for this disruption and any inconvenience it caused your team and employees during this busy season. Thank you for your patience and continued partnership. Sincerely, [Your Name/Company Name]
gpt-oss-120b: **Subject:** Update on Today’s T4 Slip Download Issue Dear HR Administrator, This morning, from 08:10 am to 09:45 am EST, employees in Canada were unable to download their T4 slips from our payroll platform. The service interruption lasted approximately 95 minutes and has now been fully resolved. Please have your employees retry downloading their T4 slips. If the download succeeds, no further action is required on your part. All payroll data, including the T4 information, remained securely stored and unchanged throughout the incident. Our team is actively monitoring the system to ensure continued stability. We understand how critical timely access to tax documents is for your team, and we sincerely apologize for any inconvenience this may have caused. Your confidence in our service is important to us, and we remain committed to delivering reliable, secure payroll solutions. If you have any questions, please contact our support team at support@payrollsaaS.com. Thank you for your understanding. Sincerely, [Your Name] Customer Success Manager

Model A better matches the requested calm, accountable, non-technical tone and stays tightly aligned to the required content without adding unnecessary specifics. Model B is also strong, but the exact outage timestamps and more formal operational phrasing feel slightly more technical and less audience-tailored for HR administrators. (Second judge pass, order swapped — scores are the average of both: Model A better matches the requested calm, accountable, non-technical customer update while clearly covering the outage, next steps, data integrity reassurance, and apology within the target length. Model B is also strong, but it is slightly more technical/formal in tone and includes exact timestamps and monitoring language that are less aligned with the brief.)

Messy shipment extraction

Pull structured data from this messy note. Return a JSON object with exactly these keys: order_id, customer_name, ship_to_city, items, total_qty, promised_ship_date, tracking_code, fragile, support_case. Text: "KiteLine Ops scribble — pls rush. ORD? it's ZX-4819-B. Customer: Marisol Venn (billing unchanged). Ship-to: 14 Quarry Mews, Dunbar, OR 97111 — city is Dunbar. Cart had: 3x 'AeroLatch mini' sku ALM-04, 1 x 'Copperfin filter' sku CFF-9, and 2x AeroLatch mini again added by phone. Ignore canceled draft item: 5x 'Nightglass'. Promised ship date moved from 2026-04-12 to Tue 2026-04-14. Tracking draft said TRK-0091-XY but final label reads TRK-0091-XY-2. Mark FRAGILE? yes, because of filter housing. Related support case: CS-77102."

Google: Gemini 3.6 Flash: ```json { "order_id": "ZX-4819-B", "customer_name": "Marisol Venn", "ship_to_city": "Dunbar", "items": [ { "name": "AeroLatch mini", "sku": "ALM-04", "quantity": 5 }, { "name": "Copperfin filter", "sku": "CFF-9", "quantity": 1 } ], "total_qty": 6, "promised_ship_date": "2026-04-14", "tracking_code": "TRK-0091-XY-2", "fragile": true, "support_case": "CS-77102" } ```
gpt-oss-120b: ```json { "order_id": "ZX-4819-B", "customer_name": "Marisol Venn", "ship_to_city": "Dunbar", "items": [ { "name": "AeroLatch mini", "sku": "ALM-04", "quantity": 5 }, { "name": "Copperfin filter", "sku": "CFF-9", "quantity": 1 } ], "total_qty": 6, "promised_ship_date": "2026-04-14", "tracking_code": "TRK-0091-XY-2", "fragile": true, "support_case": "CS-77102" } ```

Model A and Model B are identical and correctly extract all requested fields, including combining the repeated AeroLatch mini quantities, ignoring the canceled item, and using the final tracking code and updated promised ship date. Both follow the instruction to return exactly the required JSON structure and present the data clearly. (Second judge pass, order swapped — scores are the average of both: Model B and Model A are identical and correctly extract all requested fields, including combining the repeated AeroLatch mini quantities, ignoring the canceled item, and using the final tracking code and updated promised ship date. Model B and Model A also follow the required JSON-only structure cleanly and completely.)

LRU cache implementation

Implement an LRU cache in Python 3. Create a class `LRUCache` with methods `get(key: int) -> int` and `put(key: int, value: int) -> None`, initialized by `LRUCache(capacity: int)`. Both operations must run in O(1) time. When capacity is exceeded, evict the least recently used key. If `get` misses, return -1. Use your own doubly linked list + hash map; do not use `OrderedDict`. Return code only.

Google: Gemini 3.6 Flash: ```python class Node: def __init__(self, key: int = 0, value: int = 0): self.key = key self.value = value self.prev = None self.next = None class LRUCache: def __init__(self, capacity: int): self.capacity = capacity self.cache = {} # Dummy head and tail nodes to simplify list manipulation self.head = Node() self.tail = Node() self.head.next = self.tail self.tail.prev = self.head def _remove(self, node: Node) -> None: """Remove an existing node from the doubly linked list.""" prev_node = node.prev next_node = node.next prev_node.next = next_node next_node.prev = prev_node def _add_to_tail(self, node: Node) -> None: """Insert a node right before the dummy tail (most recently used).""" prev_node = self.tail.prev prev_node.next = node node.prev = prev_node node.next = self.tail self.tail.prev = node def get(self, key: int) -> int: if key in self.cache: node = self.cache[key] self._remove(node) self._add_to_tail(node) return node.value return -1 def put(self, key: int, value: int) -> None: if key in self.cache: node = self.cache[key] node.value = value self._remove(node) self._add_to_tail(node) else: if len(self.cache) >= self.capacity: # Evict the least recently used node (head.next) lru_node = self.head.next self._remove(lru_node) del self.cache[lru_node.key] new_node = Node(key, value) self.cache[key] = new_node self._add_to_tail(new_node) ```
gpt-oss-120b: ```python from __future__ import annotations from typing import Optional class _Node: __slots__ = ("key", "value", "prev", "next") def __init__(self, key: int, value: int): self.key: int = key self.value: int = value self.prev: Optional[_Node] = None self.next: Optional[_Node] = None class LRUCache: """ Least‑Recently‑Used (LRU) cache with O(1) get and put operations. """ def __init__(self, capacity: int): if capacity <= 0: raise ValueError("capacity must be > 0") self._capacity: int = capacity self._cache: dict[int, _Node] = {} # Dummy head & tail simplify insert / removal logic self._head: _Node = _Node(0, 0) # most‑recently used side self._tail: _Node = _Node(0, 0) # least‑recently used side self._head.next = self._tail self._tail.prev = self._head # ------------------------------------------------------------------ # # Internal helper methods (all O(1)) # ------------------------------------------------------------------ # def _remove(self, node: _Node) -> None: """Detach `node` from the linked list.""" prev_node = node.prev nxt_node = node.next if prev_node is not None and nxt_node is not None: prev_node.next = nxt_node nxt_node.prev = prev_node node.prev = node.next = None def _add_to_front(self, node: _Node) -> None: """Insert `node` right after the dummy head (most‑recent position).""" node.prev = self._head node.next = self._head.next self._head.next!.prev = node # type: ignore[union-attr] self._head.next = node def _move_to_front(self, node: _Node) -> None: """Mark `node` as most recently used.""" self._remove(node) self._add_to_front(node) def _evict_lru(self) -> None: """Remove the least recently used node (right before dummy tail).""" lru = self._tail.prev if lru is self._head: # cache empty; should not happen return self._remove(lru) del self._cache[lru.key] # ------------------------------------------------------------------ # # Public API # ------------------------------------------------------------------ # def get(self, key: int) -> int: """Return value for `key` if present, else -1. Marks key as recently used.""" node = self._cache.get(key) if node is None: return -1 self._move_to_front(node) return node.value def put(self, key: int, value: int) -> None: """Insert or update `key` with `value`. Evicts LRU entry if capacity exceeded.""" node = self._cache.get(key) if node: # Update existing entry and move to front node.value = value self._move_to_front(node) else: # Insert new entry if len(self._cache) >= self._capacity: self._evict_lru() new_node = _Node(key, value) self._cache[key] = new_node self._add_to_front(new_node) ```

Model A provides a valid O(1) LRU cache using a hash map plus a custom doubly linked list and is close to production-ready, aside from not guarding the edge case of zero capacity. Model B is not valid Python because it uses `self._head.next!`, so despite otherwise solid structure, it fails correctness and therefore loses. (Second judge pass, order swapped — scores are the average of both: Model A provides a valid O(1) hash map plus doubly linked list LRU implementation that largely satisfies the prompt. Model B is not runnable Python because it uses invalid syntax (`self._head.next!`) and also includes extra non-code output despite the instruction to return code only.)

Precise proofreading

Correct ONLY the grammar, spelling, and punctuation errors in the sentence below without changing wording, tone, or meaning, then on a second line list each change as 'was -> now'. Sentence: "Its been a long time since we seen there new office, and me and him was suprised how quite it was."

Google: Gemini 3.6 Flash: "It's been a long time since we saw their new office, and he and I were surprised how quiet it was." Its -> It's, seen -> saw, there -> their, me and him -> he and I, was -> were, suprised -> surprised, quite -> quiet
gpt-oss-120b: It's been a long time since we saw their new office, and he and I were surprised how quiet it was. Its -> It's, seen -> saw, there -> their, me -> I, him -> he, was -> were, suprised -> surprised, quite -> quiet

Model B better follows the requested format by placing the corrected sentence on one line and the change list on a second line without quotation marks, while still making the needed corrections. Model A is also correct, but it adds quotation marks and groups one change as a phrase replacement rather than listing each change as explicitly as Model B. (Second judge pass, order swapped — scores are the average of both: Model B is better because it provides the corrected sentence without adding unnecessary quotation marks and lists the changes clearly. Model A introduces quotation marks not present in the original and combines "me and him" into one change entry, making its change list less precise for a proofreading task.)

Proofread incident note

Correct the passage below for grammar, punctuation, and clarity without changing its meaning. Then provide 3 brief bullet points noting the main changes. "Yesterdays deploy to the invoices service were suppose to reduce timeout's, but instead it increased p95 latency from 380ms to 1.4s on only eu-west customers. We dont think the database is the cause because CPU stayed under 35%, however the new retry logic are probably amplifying traffic. By 16:40 we roll backed the worker image and latency returned near baseline, though a few requests was still queued for around 20 minutes."

Google: Gemini 3.6 Flash: **Corrected Passage:** "Yesterday's deployment to the invoices service was supposed to reduce timeouts, but instead it increased p95 latency from 380ms to 1.4s for eu-west customers only. We don't think the database is the cause because CPU usage stayed under 35%; however, the new retry logic is probably amplifying traffic. By 16:40, we rolled back the worker image and latency returned to near baseline, though a few requests were still queued for around 20 minutes." *** ### Main Changes Made: * **Subject-Verb Agreement & Tense:** Corrected verb forms throughout the passage (e.g., "were suppose" to "was supposed," "logic are" to "logic is," "roll backed" to "rolled back," and "requests was" to "requests were"). * **Punctuation & Apostrophes:** Added missing apostrophes for possessives/contractions ("Yesterday's," "don't"), removed an incorrect plural apostrophe ("timeouts"), and added a semicolon before "however" to fix a comma splice. * **Clarity & Phrasing:** Replaced the informal noun "deploy" with "deployment," added "usage" after CPU, and improved prepositional phrasing (e.g., "returned to near baseline").
gpt-oss-120b: **Revised passage** Yesterday’s deployment to the invoices service was supposed to reduce timeouts, but instead it increased the p95 latency from 380 ms to 1.4 s for EU‑West customers only. We don’t think the database is the cause because CPU stayed under 35 %, but the new retry logic is probably amplifying traffic. By 16:40 we had rolled back the worker image, and latency returned to near baseline, though a few requests were still queued for about 20 minutes. **Main changes** - Fixed possessive forms, plurals, and verb tenses (e.g., “Yesterday’s deployment,” “was supposed,” “are → is,” “rolled back,” “were still queued”). - Corrected punctuation and added missing commas for clearer sentence flow. - Clarified wording and consistency (e.g., “time‑outs” → “timeouts,” “EU‑West customers only,” “CPU stayed under 35 %”).

Model A is better because it corrects the passage cleanly while preserving meaning and provides three concise, accurate bullets about the main edits. Model B is also strong, but it introduces a tense shift with "we had rolled back" that slightly changes the timeline and its change notes are a bit less precise. (Second judge pass, order swapped — scores are the average of both: Model B is slightly better because it improves grammar and clarity while preserving meaning more precisely, especially with the more polished phrasing and cleaner final sentence. Model A is also strong, but it leaves "eu-west" uncapitalized and makes a slightly more interpretive change with "CPU usage," while Model B’s revision reads more naturally overall.)

Constraint scheduling

Four talks (A, B, C, D) fill four 1-hour slots 9,10,11,12. Constraints: A is before D; C is not first; B is immediately after A; D is not at 12. Give the ONE valid schedule as 'slot: talk' lines, then a one-line justification. If impossible, say so and explain.

Google: Gemini 3.6 Flash: 9: A 10: B 11: D 12: C Justification: Since B must immediately follow A and A must precede D (which cannot be at 12), A and B must take slots 9 and 10, leaving D for 11 and C for 12 to satisfy all conditions.
gpt-oss-120b: 9 AM – A 10 AM – B 11 AM – D 12 PM – C **Justification:** A precedes D, B follows A directly, C is not in the first slot, and D is not at 12 PM. All constraints are satisfied.

Model A and Model B both give the unique valid schedule and satisfy the constraints, but Model A is better because it follows the requested output format more closely with explicit 'slot: talk' lines and provides a slightly stronger justification for uniqueness. Model B is correct and clear, but its formatting departs a bit from the specified format. (Second judge pass, order swapped — scores are the average of both: Model A is slightly better because it gives the same correct unique schedule while matching the requested 'slot: talk' format more closely and providing a concise justification for why the solution is forced. Model B is correct, but its formatting uses AM/PM dashes rather than the specified slot-line style.)

Normalize phone numbers

In JavaScript, write a function `normalizePhones(records)` that takes an array of objects like `{name, phone}` and returns a new array where `phone` is normalized to `+1XXXXXXXXXX` for valid US numbers. Accept inputs such as `"(541) 555-0188"`, `"1-541-555-0188"`, `"541.555.0188 x22"`, stripping punctuation and any extension. If the remaining digits are 10 digits, prefix `+1`; if 11 digits starting with `1`, convert to `+1` plus the last 10 digits; otherwise set `phone` to `null`. Do not mutate the input. Return code only.

Google: Gemini 3.6 Flash: ```javascript function normalizePhones(records) { return records.map(record => { let normalizedPhone = null; if (record.phone !== null && record.phone !== undefined) { const mainPart = String(record.phone).split(/x|ext|#/i)[0]; const digits = mainPart.replace(/\D/g, ''); if (digits.length === 10) { normalizedPhone = `+1${digits}`; } else if (digits.length === 11 && digits.startsWith('1')) { normalizedPhone = `+1${digits.slice(1)}`; } } return { ...record, phone: normalizedPhone }; }); } ```
gpt-oss-120b: ```javascript function normalizePhones(records) { return records.map(record => { const { phone } = record; let normalized = null; if (typeof phone === 'string') { // Remove any extension (e.g., "x22", "X123", "ext. 45") const noExt = phone.split(/[xX]/)[0]; // Keep only digits const digits = noExt.replace(/\D/g, ''); if (digits.length === 10) { normalized = '+1' + digits; } else if (digits.length === 11 && digits[0] === '1') { normalized = '+1' + digits.slice(1); } } // Return a shallow copy with the normalized phone return { ...record, phone: normalized }; }); } ```

Model A better matches the prompt because it strips extensions more robustly, handling patterns like "x", "ext", and "#", while still avoiding input mutation and correctly normalizing 10- and 11-digit US numbers. Model B is mostly correct but only removes extensions via "x"/"X", so inputs like "ext. 45" are mishandled despite its comment claiming otherwise. (Second judge pass, order swapped — scores are the average of both: Model A better handles extension stripping by recognizing more common markers like "ext" and "#", while still preserving immutability and applying the required 10-digit and 11-digit normalization rules. Model B is solid, but its extension handling is narrower and can mis-handle inputs containing an "x" earlier than an actual extension marker.)

Constrained meeting summary

Summarize this meeting in exactly 4 bullet points. Each bullet must start with one of these tags in order: `DECISION:`, `RISK:`, `OWNER:`, `DATE:`. Each bullet must be 8-14 words. Do not use the words "meeting", "discussed", or "issue". Notes: "The team chose Northwind Studio to redesign the tenant onboarding flow. Biggest concern: legal review could delay launch if the updated consent screen wording isn't approved quickly. Priya will send revised copy to legal and to Northwind by Thursday afternoon. Target launch remains 18 September 2026, assuming approval lands next week."

Google: Gemini 3.6 Flash: * DECISION: The team selected Northwind Studio to redesign the tenant onboarding workflow. * RISK: Delayed legal review of updated consent screen wording might postpone the launch. * OWNER: Priya will deliver revised consent copy to legal and Northwind by Thursday. * DATE: Target launch remains 18 September 2026, assuming legal approval arrives next week.
gpt-oss-120b: - DECISION: Select Northwind Studio to redesign the tenant onboarding flow. - RISK: Legal review may delay launch if consent wording not approved promptly. - OWNER: Priya sends revised copy to legal and Northwind by Thursday. - DATE: Target launch set for 18 September 2026, pending approval.

Model B better satisfies the strict length constraint: all four bullets appear to stay within 8-14 words, while Model A’s DECISION bullet exceeds the limit. Both are accurate and avoid the banned words, but Model B is more compliant overall and remains clear and concise. (Second judge pass, order swapped — scores are the average of both: Model B better satisfies the strict constraints: all four bullets use the required tags in order, avoid banned words, and each line stays within 8-14 words. Model A is accurate and well-written, but every bullet exceeds the 14-word limit, so its instruction following is weaker.)

SQL window query

Given Postgres tables `orders(id, customer_id, created_at, total_cents)` and `customers(id, name)`, write ONE SQL query returning each customer's name and their SECOND-largest order total (in dollars, 2 decimals), excluding customers with fewer than two orders. Order by that value descending. Return ONLY the SQL.

Google: Gemini 3.6 Flash: ```sql WITH ranked_orders AS ( SELECT c.name, o.total_cents, ROW_NUMBER() OVER (PARTITION BY o.customer_id ORDER BY o.total_cents DESC) AS rn FROM customers c JOIN orders o ON c.id = o.customer_id ) SELECT name, ROUND(total_cents / 100.0, 2) AS second_largest_order_total FROM ranked_orders WHERE rn = 2 ORDER BY second_largest_order_total DESC; ```
gpt-oss-120b: ```sql SELECT c.name, round(o.total_cents / 100.0, 2) AS second_largest_total FROM ( SELECT customer_id, total_cents, row_number() OVER (PARTITION BY customer_id ORDER BY total_cents DESC) AS rn FROM orders ) o JOIN customers c ON c.id = o.customer_id WHERE o.rn = 2 ORDER BY second_largest_total DESC; ```

Model A and Model B both correctly use ROW_NUMBER() partitioned by customer to select the second-highest order, exclude customers with fewer than two orders by filtering to rn = 2, convert cents to dollars with two decimals, and order descending by that value. Model A is slightly more explicit with the CTE, while Model B is slightly more concise, but neither has a meaningful quality advantage for this task. (Second judge pass, order swapped — scores are the average of both: Model B and Model A both correctly use ROW_NUMBER() partitioned by customer to select the second-highest order per customer, exclude customers with fewer than two orders by filtering to rn = 2, convert cents to dollars with two decimals, and order descending by that value. Model B is slightly more direct while Model A is slightly more verbose, but both fully satisfy the prompt with equally strong quality.)

LRU cache

Implement a class `LRUCache<K, V>` in TypeScript with a fixed capacity set in the constructor, and O(1) `get(key): V | undefined` and `set(key, value): void`. Accessing or updating a key must mark it most-recently-used; inserting beyond capacity must evict the least-recently-used entry. Return ONLY the class, no prose.

Google: Gemini 3.6 Flash: ```typescript class LRUCache<K, V> { private readonly capacity: number; private readonly cache: Map<K, V>; constructor(capacity: number) { if (capacity <= 0) { throw new Error("Capacity must be greater than 0"); } this.capacity = capacity; this.cache = new Map<K, V>(); } get(key: K): V | undefined { if (!this.cache.has(key)) { return undefined; } const value = this.cache.get(key)!; // Move to end to mark as most-recently-used this.cache.delete(key); this.cache.set(key, value); return value; } set(key: K, value: V): void { if (this.cache.has(key)) { this.cache.delete(key); } else if (this.cache.size >= this.capacity) { // Delete the first key (least-recently-used) const lruKey = this.cache.keys().next().value; if (lruKey !== undefined) { this.cache.delete(lruKey); } } this.cache.set(key, value); } } ```
gpt-oss-120b: ```ts class LRUCache<K, V> { private readonly capacity: number; private readonly cache: Map<K, V>; constructor(capacity: number) { if (capacity < 1) throw new Error('Capacity must be greater than 0'); this.capacity = capacity; this.cache = new Map(); } get(key: K): V | undefined { if (!this.cache.has(key)) return undefined; const value = this.cache.get(key)!; // Move to most-recently-used position this.cache.delete(key); this.cache.set(key, value); return value; } set(key: K, value: V): void { if (this.cache.has(key)) { this.cache.delete(key); } else if (this.cache.size >= this.capacity) { // Evict least-recently-used (first inserted) key const lruKey = this.cache.keys().next().value; this.cache.delete(lruKey); } this.cache.set(key, value); } } ```

Model A and Model B both implement the required API and LRU behavior in O(1) average time using Map insertion order, but Model A is slightly stronger because it avoids deleting an undefined eviction key by guarding the iterator result. Model A also uses an explicit generic type on the Map construction, making the TypeScript typing a bit cleaner and more robust. (Second judge pass, order swapped — scores are the average of both: Model A is slightly better because it is equally efficient and compliant but handles the eviction key more defensively, avoiding a potential TypeScript issue around deleting an undefined key. Model B is otherwise solid and correct in typical usage, but Model A is a bit more robust.)

Unit-aware math

A pump moves 3.5 liters every 8 seconds. A tank holds 0.9 cubic meters. Starting empty, how long to fill it, in minutes and seconds (mm:ss), rounded to the nearest second? Show the key steps, then give the final answer on its own line.

Google: Gemini 3.6 Flash: **Key Steps:** 1. **Convert the tank volume to liters:** $1\text{ m}^3 = 1,000\text{ liters}$ $0.9\text{ m}^3 \times 1,000 = 900\text{ liters}$ 2. **Calculate the total time in seconds:** $\text{Total seconds} = \frac{900\text{ liters}}{3.5\text{ liters}} \times 8\text{ seconds}$ $\text{Total seconds} \approx 2057.14\text{ seconds}$ 3. **Convert total seconds into minutes and seconds:** $\text{Minutes} = \lfloor \frac{2057.14}{60} \rfloor = 34\text{ minutes}$ $\text{Remaining seconds} = 2057.14 - (34 \times 60) = 17.14\text{ seconds}$ 4. **Round to the nearest second:** $17.14\text{ seconds} \approx 17\text{ seconds}$ 34:17
gpt-oss-120b: **Key steps** 1. Convert tank volume to liters: \(0.9\ \text{m}^3 = 0.9 \times 1000 \text{ L} = 900\text{ L}\). 2. Find the pump’s flow rate: \(\displaystyle \frac{3.5\ \text{L}}{8\ \text{s}} = 0.4375\ \text{L/s}\). 3. Compute the time to deliver 900 L: \[ t = \frac{900\ \text{L}}{0.4375\ \text{L/s}} = \frac{900}{0.4375}\ \text{s} = \frac{900}{7/16}\ \text{s} = 900 \times \frac{16}{7} = \frac{14400}{7}\ \text{s} \approx 2057.14\ \text{s}. \] 4. Convert seconds to minutes + seconds: \[ 2057.14\ \text{s} = 34\ \text{min} + 17.14\ \text{s} \] Rounded to the nearest second → 17 s. **Final answer** 34:17

Model B is slightly better because it is fully unit-aware throughout, explicitly computes the flow rate, and presents the derivation more cleanly while still following the requested format. Model A is also correct and well-structured, but its time formula is a bit less explicit about units and setup. (Second judge pass, order swapped — scores are the average of both: Model B and Model A both arrive at the correct result and show key steps, but Model B is slightly better because it explicitly computes the flow rate with units before solving, making the reasoning clearer and more unit-aware. Model A is also correct and concise, but its intermediate setup is a bit less explicit.)

Matchup powered by OpenRouter.