In November 2021, at an urban training site on Fort Campbell, Kentucky, DARPA closed out its Offensive Swarm-Enabled Tactics program by putting more than 100 air and ground robots under the direction of one person. The program manager, Timothy Chung, had been blunt about the goal. “I don’t want 100-250 Soldiers or Marines running around with joysticks in their hands, heads down,” he told the Army’s news service, describing instead a single operator working alongside a swarm commander through tablet and virtual-reality interfaces.
The headline number was the swarm size. The more useful result came from Oregon State University researchers who instrumented the swarm commanders and published their findings in the journal Field Robotics. Commanders reported their workload every 10 minutes and, in the final exercises, wore physiological sensors. According to Julie Adams, the Oregon State professor who led the work, the workload estimate “did cross the overload threshold frequently, but just for a few minutes at a time,” and the missions were still completed.
One person supervised a hundred machines because those machines rarely needed that person, and when they did, the demand came in short spikes that the commander could absorb. The capacity of a one-to-many system is set by how often the machines ask for a human, how long each request takes to resolve, and how hard the decision is. Vehicles per console is an output of those quantities. Programs that treat it as the requirement are measuring the wrong thing.
The arithmetic of attention was worked out years ago
Human-factors researchers formalized this problem long before the current generation of programs. In a 2007 paper in IEEE Transactions on Robotics, Jacob Crandall and Mary Cummings, then both at MIT, broke single-operator control of multiple robots into measurable parts. Neglect time is how long a robot can be ignored before its performance falls below an acceptable level. Interaction time is how long the operator needs to orient to that robot’s situation, decide what to do and enter the command. On top of those sit the costs of dividing attention: the time spent choosing which robot to service next, and the time a robot waits in a degraded state because the operator is busy elsewhere or has lost track of it.
Their experiment had participants run teams of two, four, six and eight simulated robots through a timed retrieval task. Output rose with team size up to six robots and then leveled off. Teams of six and eight also lost significantly more robots than teams of two and four, and operators grew worse at picking the robot that most needed attention as the team grew. The authors put the most effective team size at between four and six, and noted that the answer depended on how much a lost robot was worth relative to the objective. The robots were simulated, and the authors cautioned that the specific numbers would not transfer to real ones.
A second point from the literature matters as much. A 2016 study in Frontiers in Psychology, conducted with experienced unmanned-aircraft operators, concluded that one operator could supervise the health and status of up to 15 aircraft efficiently but could not control the mission and payload of more than three at the level of automation examined. The same paper summarizes earlier Cummings results in which operators handled four to five vehicles when the automation required their consent and eight to 12 when it acted unless they objected. A ratio quoted without the task and the rules for intervention attached is close to meaningless.
Current programs still describe themselves in airframes
The Air Force’s Collaborative Combat Aircraft program has been sized by ratio from the start. The Congressional Research Service traces the notional figure of 1,000 aircraft to a 2023 assumption of two for each of roughly 500 advanced crewed fighters. On June 17, 2026, the service awarded production contracts to General Atomics and Anduril for a first increment of at least 150 aircraft.
What a pilot will actually be asked to do with them is less settled. The fiscal 2026 budget request included about $12.2 million for 142 tablets and cabling so that F-22 pilots can direct the drones. In October 2025 an F-22 pilot commanded a single MQ-20 Avenger from the cockpit over the Nevada Test and Training Range, in a company-funded demonstration announced by General Atomics. In January 2026 the Navy reported that F-35 pilots in its Joint Simulation Environment used touch-screen tablets to control “multiple” aircraft, without saying how many or what it cost the pilots in attention.
People close to the work have been candid about that cost. Michael Atwood of General Atomics said in 2024 that when he flew with a tablet, “it was really hard to fly the airplane, let alone the weapon system of my primary airplane, and spatially and temporally think about this other thing.” John Clark, then head of Lockheed Martin’s Skunk Works, called the tablet possibly the fastest way to begin experimenting and said it “may not be the end state.” Air Force testers appear to share the concern. A team at Eglin Air Force Base recently flew an F-15E and an F-16 with two autonomous XQ-58s, and the stated purpose was to understand pilot workload and situational awareness, as Aerospace America reported in July. The first large-force exercise with the new aircraft, Emerald Flag, is planned for December, and the Experimental Operations Unit’s director of operations has said the lessons about pilot trust will be as valuable as the technical ones.
The ground and maritime services are earlier on the same path. An Army colonel responsible for maneuver requirements said in October 2024 that the service in some cases had “two Soldiers to one robot,” and a cavalry troop at the National Training Center that September needed three control vehicles to run four robotic combat vehicles. Lt. Gen. Robert Rasch, who oversees the Army’s human-machine integrated formations effort, described the interface problem in April 2025: “Today, everything has its own controller.” The Navy stood up Unmanned Surface Vessel Squadron Three in May 2024 and created a robotics warfare specialist rating to crew it, but its announcements describe the squadron by its boats and its 400 sailors and say nothing about how many craft one sailor is expected to supervise.
In combat, the ratio still runs the other way
Fielded systems show how far the ambition is from practice. A 2021 Center for Strategic and International Studies brief found that keeping one MQ-1 or MQ-9 combat line airborne around the clock takes 10 pilots and 10 sensor operators, inside a mission control element of 49 people and a forward launch and recovery element of 59.
Ukraine’s drone units are organized the same way at smaller scale. A first-person-view strike team typically needs three to six people. One unit using swarming software from the Ukrainian firm Swarmer typically flew three drones per mission, with other units reported to have flown up to eight, according to a September 2025 summary of Wall Street Journal reporting. The company’s founder says a mission that once took nine people now takes three. That is a vendor’s claim, and even taken at face value it describes three people supervising a handful of aircraft. A CSIS assessment in March 2025 concluded that fully realized swarms “have yet to be developed” there.
These accounts fit the human-factors model. Software that allocates tasks among aircraft lengthens neglect time, while planning, strike authorization and recovery from failures remain human work and keep several people on each mission.
Write the requirement in interventions
Requirements and test reports should change first. A program that promises one-to-many control should state, for a defined mission, how often each vehicle is expected to need a human, the median and worst-case time to resolve a request, and how long a vehicle can wait before the mission degrades. Tests should report those figures under jamming and partial loss of link, with the operator carrying a realistic primary task such as flying a fighter or commanding a platoon. The Oregon State team showed that this can be instrumented in the field. Emerald Flag is the next obvious place to do it, and the Air Force should publish what it finds.
Interfaces should then be judged by how much they shorten interaction time and how well they direct attention. An operator needs to tell at a glance whether a vehicle is waiting for permission, has lost its link or is unable to finish its task, because each calls for a different response and a different urgency. Manning should follow from the measured peak in intervention demand, and the Army’s control-vehicle experience is a caution against optimistic averages. A formation whose operators are saturated for a few minutes at a time in a permissive test may be saturated continuously against an enemy who is trying to cause exactly that.
This analysis draws on the public sources linked in the text. Send corrections to [email protected].


