DeskForge Corpus Improves Computer-Use Agent Accuracy
A new corpus of annotated desktop observations, DeskForge, enables computer-use agents to achieve higher accuracy in complex desktop scenes.

Computer-use agents require reliable action target grounding in complex desktop scenes. This challenge is addressed by the DeskForge corpus, which provides a dataset for training computer-use agents to better navigate these scenarios. According to arXiv cs.CV, the existing training data often lack dense annotations and vary little in a controlled manner, making it difficult to achieve accurate results. DeskForge varies multiple factors, including application states, content, window layout, and resolution, allowing for a more comprehensive and realistic training environment. The DeskForge environment provides a corpus of 1.2 million annotated desktop observations containing 159.7 million element instances. This large dataset enables the training of more accurate computer-use agents, which can better handle complex desktop scenes. Four vision-language models were fine-tuned on 200,000 grounding examples drawn from the DeskForge corpus, resulting in improved performance. One model, Qwen3.5-4B, achieved an accuracy increase of 11.51 percentage points on two specific benchmarks. The DeskForge corpus has significant implications for the development of computer-use agents. With this dataset, researchers can train agents that can better handle complex desktop scenes, which is essential for applications such as automated data entry and desktop-based tasks. The improved accuracy of these agents can lead to increased efficiency and reduced errors in these applications. In conclusion, the DeskForge corpus is an important advancement in the field of computer-use agents. The large and diverse dataset provides a platform for training agents that can better handle complex desktop scenes, leading to improved performance and increased efficiency.
Sources
Related stories

Model Auditing: Faithfulness of Post-hoc Explanations Investigated
A study published on arXiv cs.CV investigates the faithfulness of post-hoc explanations for a pedestrian detection model across different domains.

edithero benchmark for long-horizon 3d editing
The EditHero benchmark is introduced to evaluate and improve the reliability of long-horizon, part-level 3D editing.

ComfyUI v0.39.0 Release Brings New Features and Improvements
The ComfyUI v0.39.0 release includes dynamic group widgets, new model nodes, and resolved minimax vae offload issue.

Predictive Analytics Meets AI
The predictive analytics landscape has evolved significantly in recent years, with AI-powered analytics driving growth and innovation across various industries.