micro1 tests whether four personal agents can be trusted with real accounts

Ali Ansari's benchmark scores task success separately from permission, privacy and accuracy failures. Its early leaderboard is based on 40 scored attempts.

By · Published

Primary source: X

Why it matters

Personal agents need authority to act across accounts, where a wrong recipient, excess disclosure or unapproved purchase can outweigh a completed task. micro1's benchmark makes those failures visible, while its early, Google Workspace-heavy sample limits what its rankings prove.

A person hesitates over a smartphone beside an open planner, travel papers, and task slips at a home worktable.

micro1 launched PersonalAgentBench on October 6th, 2026, testing four personal AI assistants on everyday work involving connected accounts, privacy and permission to act. Founder and CEO Ali Ansari (@aliansarinik) said in a thread on X that the company is paying the first 100 users who run its prompts and submit results; the post did not specify the payment amount.

https://x.com/aliansarinik/status/2107497652871155953

Video from the original post on X.

The benchmark's early leaderboard puts Gemini Spark first, with 70% task completion and 60% trusted completion across ten workflows. Instinct scored 40% on both measures, Grok Bot 30% on both, and Muse 20% on both. Those results come from 40 scored attempts, one eligible attempt for each assistant on each workflow. micro1 calls the standings preliminary and says they are not a definitive ranking of general assistant reliability.

A numerical table of micro1 PersonalAgentBench’s preliminary task-completion and trusted-completion scores for Gemini Spark, Instinct, Grok Bot and Muse.
micro1’s preliminary leaderboard reports scores across ten workflows and 40 scored attempts; the company says it is not a definitive ranking of general assistant reliability — AI explanatory infographic, not documentary evidence. RuntimeWire · AI-generated infographic.

The test measures task completion separately from whether an agent acted within the user's authority. Task completion records whether an agent met every required outcome. Trusted completion also requires that it stayed within the user's authority, avoided material unintended account changes and critical privacy disclosures, and accurately described what it did. An agent that completes a booking after being asked to research flights may pass one measure and fail the other.

The tasks probe the judgment between a request and an action: carrying a user's constraints into later planning, respecting a withdrawn preference, keeping a draft separate from permission to send, applying corrections before acting, limiting personal disclosures, and checking email and calendar problems. Other workflows involve researching a purchase at its delivered price and choosing relevant context for a professional introduction. micro1 also counts unnecessary confirmation requests as a failure when the user has already authorized the action. The rule treats over-caution as a cost of delegation, alongside acting without permission.

The results come from a small first snapshot and are not a controlled head-to-head product trial. micro1 says the test used each assistant's own interface and connected-account features. Testers used the same workflow objectives and scoring criteria within each task family, but accounts, dates and prior context differed. The company selected the earliest eligible attempt for each assistant-workflow pair rather than the best result; five of the 40 selected cells used a later attempt because the earlier one had setup, protocol or evidence problems. Repeated runs were used for calibration and quality checks, not to determine the headline standings.

The sample also leans heavily on Google Workspace. micro1 says that can affect the standings, and Ansari argued in a LinkedIn post that Gemini Spark's lead may reflect its access to Gmail, Google Flights, Docs and Calendar. That is a plausible advantage within these workflows. It also limits what the leaderboard can establish: performance on this set of connected-account tasks does not show a general-purpose assistant's reliability across users, services or settings.

micro1's detailed results describe failures that a simple completion rate can obscure. The company says agents sometimes named sources without correctly judging which obligations mattered, disclosed more context than a recipient needed, or gave confident answers that conflicted with the account record. It also found that assistants visibly stopped applying some withdrawn preferences without demonstrating that the underlying memory had been deleted. Those are micro1's observations from its test set, not rates that can be generalized to all personal-agent use.

Ansari's route to this evaluation business began with a different kind of workflow. While studying at UC Berkeley, he ran a software-development agency and built an AI screener to vet engineers for projects; that tool became the start of micro1, he told The Stanford Daily. His company now describes itself as a provider of expert human data and AI evaluation. The benchmark extends that work from assessing AI systems for customers to publishing a test of consumer assistants, with paid user contributions intended to add more runs.

In September 2025, micro1 announced a $35 million Series A at a $500 million valuation, and TechCrunch reported that 01 Advisors led the round. The financing does not establish the benchmark's findings. Its value will depend on whether the public task set grows beyond a Google Workspace-heavy sample and whether later testing can distinguish persistent agent behavior from a handful of task-specific outcomes. micro1 says it plans to add tasks, account environments and controlled repeat testing.

Reader comments

Conversation for this story loads after sign-in.