Most AI assistants still answer in a chat box. A newer class of model tries to do the thing you asked instead, moving a cursor, tapping a phone screen, filling in a form. On 20 August Alibaba released its entry, Qwen-UI-Agent, and the numbers it published are the kind that get attention.

The agent is built to operate across four settings: mobile phones, desktop computers, web browsers, and deep search across the open web. On MobileWorld, a benchmark for driving real phone apps, Alibaba reports a score of 82.1 percent, which it says beats OpenAI's GPT-5.6 Sol by 12 points and Claude Opus 4.8 by nearly 15. On a harder test run against physical handsets, the company claims 92.2 percent, and on a daily-tasks Android evaluation it puts the figure at 97.5 percent. Desktop control, measured on OSWorld-Verified, comes in at 79.5 percent, and web navigation on WebArena at 73.6 percent.

What the model is meant to do

The appeal of a GUI agent is that it does not need a special interface. It looks at the screen the way a person does, finds the button, and presses it, which in principle lets it use any app rather than only the ones with a developer-friendly connection. Alibaba says it trained the system on real interactions gathered from more than 100 mobile devices and over 150 applications, and released a technical report describing the approach. Independent developers have noted that the weights can run on ordinary hardware, which puts a capable screen-controlling agent within reach of people outside the big labs, much as the recent wave of open Chinese coding models has done for software work.

Read the scoreboard carefully

A note of caution belongs here. These are Alibaba's own reported results, on benchmarks the company selected, and no outside group has yet reproduced them. Vendor numbers tend to flatter the vendor. It is also worth remembering that screen agents remain brittle in the messy real world: a moved button, a surprise pop-up, or a login screen can derail a run that looked flawless on a benchmark. High marks on a curated test are a promise, not a guarantee, and the gap between the two is where most of these systems still live.

Even discounted, the direction is clear enough. The contest among the leading labs is shifting from who can hold the best conversation to who can most reliably get something done on your behalf, and a Chinese model claiming the lead on screen control, in the open, is a real marker of where that race has reached. The honest test will come when someone other than Alibaba runs the same tasks and reports what they find.

Sources

  1. i. tongyi-mai.github.io
  2. ii. arxiv.org
  3. iii. www.martincid.com
  4. iv. digitalphablet.com

Commentarii · 0

Add · a · Comment