A digital worker has to see the screen, understand the task, use the tools, remember the goal, and keep going when the easy path breaks.
It also has to be affordable enough to put to work.
That combination is the research agenda at GrupaAI. We are building toward expert digital humanoid workers that operate across computers, phones, tablets, and watches, using the apps and browsers the work requires. Our ambition is capability and accessibility together.
Our open roles are a map of the problems that stand in the way. Seven disciplines. One worker. The quality of the final experience depends on all of them.
01 / Make useful computation radically cheaper
Inference is not an isolated benchmark when a worker is running. Its cost and latency shape how often the system can perceive, think, verify, and recover.
Our inference and kernels work spans custom CUDA, Triton, and ROCm kernels; quantization; speculative and parallel decoding; mixture-of-experts routing; and serving-runtime design.
The challenge is to improve the economics of the actual workload without quietly weakening the worker. A speedup that vanishes outside a narrow benchmark is not enough. Neither is a cost reduction that creates more failed tasks.
If you have personally made a real kernel, model, or system substantially faster or cheaper, we want to understand the measurement and the trade-offs. Show the workload. Show the result. Show what did not improve.
02 / Coordinate agents across devices
An agent is easy to admire when the interface is predictable. The real problem starts when the workflow changes halfway through or moves to another device.
An element moves. A session expires. A form produces an unexpected state. The result of the last action is ambiguous. The worker must decide whether to retry, inspect, change its approach, or ask a person.
Our agents and multi-agents work connects multi-device control, shared context, and permitted integrations. That includes concurrent and parallel use of computers, phones, tablets, watches, and browsers, as well as remote device use with secure access and persistent sessions. We are interested in reliable execution across authenticated systems, including systems without a convenient API, and in teams of agents that delegate, coordinate, and recover together.
The evidence we care about is not just an elegant demo. It is what happened over repeated runs, which failures remained, and what the system did to avoid making them worse.
03 / Build the product someone will leave running
The desktop application is where abstract capability becomes a responsibility someone is willing to delegate.
That requires real operating-system depth. It also requires careful product judgment: comprehensible permissions, clear status, recoverable failures, and an obvious way to pause or redirect the worker.
The desktop worker role brings perception, hearing, voice, reasoning, and action into a coherent loop. Local performance, security, and the trustworthiness of the interface are part of the same problem.
A product someone keeps using is evidence of something a prototype cannot establish. We want to meet engineers who have shipped software that became part of a person's day, and who can explain the difficult decisions behind that dependability.
04 / Give the worker continuous perception
A screen is not a static document. It changes as the worker acts, and multiple devices may change at once.
Streaming vision must make those changes understandable: the page layout, the element a worker can interact with, the text that matters, and the evidence that an action succeeded.
Our vision work includes screen and video understanding, UI grounding, layout parsing, element detection, and OCR. Efficiency matters throughout. Distillation, caching, and resolution scheduling are possible tools, but correctness decides whether they are useful.
The goal is not to process the most pixels. It is to perceive enough, at the right time, to act responsibly.
Show us perception systems you have made work under real latency and resource constraints, including the cases where confidence should be low.
05 / Make the fleet as dependable as the worker
One worker and a large fleet are different systems problems.
Concurrent missions compete for computation and device sessions. Remote connections drop. Machines become unavailable. Workloads vary. Waiting and active execution have different needs. A low headline infrastructure price means little without utilization and recovery.
Our serving and fleet work covers batching, caching, scheduling, resource allocation, and failover across different kinds of compute. The ambition includes very large fleets, but scale targets are not a claim about the number of workers operating today.
We want engineers who have run systems, not only designed them. Bring cost curves, capacity limits, incident lessons, or a concrete example of what changed after production contradicted the plan.
06 / Give the worker hearing and a voice
Listening and speaking add another demanding real-time system.
A useful conversation requires turn-taking, interruption handling, echo cancellation, and coordination between speech and action. A worker should not keep talking because its pipeline cannot recognize that a person has interrupted it.
Our audio direction includes low-latency streaming speech recognition and synthesis, full-duplex pipelines, and a research ambition for sub-200-millisecond responsiveness. Actual latency depends on the pipeline and measurement; the target is not an achieved product guarantee.
We are interested in people who have fought these problems in functioning systems. A measured, imperfect pipeline with a clear account of its failure modes is more informative than a perfect recording.
07 / Make a smaller model excellent at a real job
The largest model is not necessarily the right model for every moment of a worker's day.
Post-training and distillation explore how to make specialized models useful at lower cost. The work includes execution-based evaluations, reward design, task data, and learning from authorized feedback.
The question is not whether a small model can imitate a large model's style. It is whether it can perform the job accurately, recognize its limits, and fit into a dependable system.
Any learning loop must respect the data permissions and confidentiality of the work. Running a task is not blanket permission to turn its contents into training material.
Show us models you have trained or distilled, the evaluations that made you believe they improved, and the cases that still required a different approach.
The hardest problem lives between the disciplines
Our operating-cost research target is below 0.3 cents per worker-hour. It is not a current price or achieved benchmark. It is a constraint that requires progress across the stack.
A faster model changes the design of the agent. Better perception changes which recovery strategies are possible. Better evaluations change what should be trained. Better product design changes how people can provide direction and oversight.
We need people who can own a difficult layer and remain curious about the whole worker. Local brilliance matters. So does the ability to find the failure that sits just outside your component.
Bring evidence, and a problem you cannot leave alone
GrupaAI's vision is larger than any one of these roles: a worker, a workforce, and ultimately an operating system for autonomous organizations.
We are hiring across the Bay Area to make the next part of that vision real. The common requirement is evidence of exceptional work. A repository, a benchmark with a reproducible method, a shipped product, a paper, or a system you operated can all tell that story.
Tell us what you personally built. Explain why it was hard. Tell us what failed and what you changed.
There is an extraordinary amount left to invent. We are looking for the people who want to own it.