A person preparing a customer proposal rarely works inside one neat sequence. They compare a document with a spreadsheet. A colleague sends a correction. A price changes. The file they were about to share is still uploading.
The task is to deliver the right proposal. Everything on the screen matters only in relation to that outcome.
This ordinary scene contains a difficult problem for a digital worker: the world can change while the worker is deciding what to do. A correct understanding of the screen a moment ago may no longer support the next action.
As we build at GrupaAI, we keep returning to a conviction: a worker should stay in contact with the work while it is happening.
The screen is part of the workplace
Software presents much of its working state visually. A document has an active version. An upload has a status. A form has a warning. A spreadsheet has a selected sheet, a filter, and sometimes a row that changes the meaning of every total beneath it.
Those details can determine whether an action is useful. They also explain why a plausible plan can fail when it meets an actual application.
We see visual understanding as an important foundation for a digital humanoid worker. The term describes our ambition for software that can operate within human working environments. It does not mean that the software is a person, or that recognizing a screen gives it human judgment.
The worker still needs a goal, context, permission, and a way to establish whether it has succeeded. A screen can ground a decision. It cannot supply all of those things by itself.
That distinction matters to how we interpret our research direction. Vision is a powerful place to begin. Its sufficiency for general work remains a question to investigate.
Stay present while the work changes
Return to the proposal. The worker has found the requested price and is preparing the document when a correction appears in the source. The old price was read accurately. Using it now would still be wrong.
We are exploring how observation, reasoning, and action can remain more closely connected, with useful observation continuing while other work proceeds. The aim is to shorten the distance between what the worker believes and what is actually happening.
There are obvious traps. More visual input can consume more resources. An animation can attract attention without changing the task. A new image does not guarantee that the important change was understood. Constantly reconsidering a decision can become another way of never finishing it.
The useful question is whether continuity helps the worker respond to meaningful change and finish correctly. A smoother recording of the same mistake would not be progress.
We want to evaluate that distinction through the result: was the proposal accurate, was the right file delivered, and how much intervention did it require?
Familiar work should become easier
A capable colleague does not approach every familiar application as though encountering it for the first time. Experience gives them a starting point. They also know when the starting point needs to change.
We think digital workers should earn a similar advantage from prior work. Familiarity should reduce repeated discovery, leaving more attention for the unfamiliar or consequential parts of a task.
That is a hypothesis about useful learning, not permission to repeat an old sequence blindly. A different account, a changed interface, or a new goal can make yesterday's successful approach unsuitable today. Private context must also stay within the boundaries that govern its use.
The ambition is accumulated competence: a worker that becomes less wasteful without becoming less observant. We should be able to test that ambition on new work, including situations in which experience is misleading.
One goal across a changing workplace
The proposal might begin in a browser and end in a desktop application. A person may review it on a phone. Some work may happen on local machines, some on authorized remote devices, while other missions continue concurrently.
Our multi-device direction follows that reality. It also spans multiple harnesses: the working environments that connect a model to its tools and the state of its task. We want the goal to remain coherent as the working surface changes. The destination should survive a change of application, screen, or execution environment.
This is a design direction, not a claim that every device and application is already supported. Each added environment creates new questions about access, context, reliability, and control. Breadth only becomes useful when a worker can complete the handoff without losing what the person asked it to do.
A visible journey should help the person understand those handoffs. It should also let them interrupt, correct, and resume the work without reconstructing the mission from the beginning.
Economics must follow completed work
Our long-term ambition is digital labor that costs no more than one-tenth as much as comparable, highly capable human work, even in the lowest-cost labor markets. That is an ambition, not an established result.
The comparison has to include equivalent quality and the full cost of delivering it: operation, supervision, recovery, and the attempts that fail. Cheap activity becomes expensive when someone must inspect and repair every result.
For the proposal, the meaningful unit is an accurate proposal delivered within the agreed boundaries. Counting only the moments when the worker was thinking would leave out much of the work that makes it dependable.
This is why perception, continuity, and economics belong in the same conversation. Better contact with the environment may reduce mistakes and repeated effort. We need to demonstrate when it does, and recognize when it adds complexity without enough benefit.
The standard we are building toward is concrete: understand the current situation, take an appropriate action, notice what changes, and carry the goal through to a result someone can use.