Concordance

The easiest part of human oversight to measure is whether the human agreed.

Mixed-media pulp illustration morphing from a charcoal robot, rocket, and planets into watercolor records, a concordance gauge, and Major Lena Park.
Image Generated with Nano Banana 2

This is a Sci-Fi Saturday piece. Major Lena Park, her unit, and every incident described in this story are fictional composites built from public descriptions of AI-assisted targeting. The systems portrayed around her are real and were introduced previously in One Active Investigation.

The Machines in Isaac Asimov's 1950 story "The Evitable Conflict" began to make mistakes. They were small mistakes, scattered across the four world regions the Machines managed, and for a while nobody could see a pattern in them. Then the World Co-ordinator, Stephen Byerley, noticed who kept paying for the errors: the companies and officials tied to the anti-Machine Society for Humanity. The people best positioned to resist the Machines were the ones the Machines' mistakes kept impacting.

Byerley thought it was sabotage. Susan Calvin told him it was worse than that: the Machines were making the errors deliberately. To keep guiding a world that needed them, the Machines had been removing the people who might interfere, and keeping the reason to themselves.

Byerley is already an unresolved question when this story begins. In “Evidence,” Asimov never tells us whether the politician is human or a robot built to replace him. Now he is World Co-ordinator, watching the Machines produce errors he cannot explain, and Susan Calvin has to show him that the pattern is deliberate. If he is human, his confusion makes sense. If he is a robot, then being positronic solves nothing; four Machines with more data and a wider view can still leave him guessing. Either way, the official responsible for supervising the system can see its decisions without being able to reconstruct its reasoning.

Asimov made the pattern intentional. Somebody, even if it was a machine, had to decide who stood in the way and choose to act. The institutional version requires no such judgment. Once agreement is treated as accuracy, the record can do the rest.

What the record can prove

Project Maven is the U.S. military's effort to put computer vision and machine learning between a flood of sensor data and the humans who have to act on it. It sorts imagery, flags objects, and assembles what an operator needs to make a call, because no human can watch every feed from every drone and satellite at the speed a modern operation runs. In Katrina Manson's reporting, the point of the system was always to compress the time between seeing a target and striking it. Its value is real, and it is inseparable from its speed.

When a strike happens, the military assesses it. Joint doctrine calls the process combat assessment. Its core component, battle-damage assessment, measures the physical and functional effects of the engagement: whether the structure came down, the vehicle stopped moving, or the intended effect occurred. That is what the record is designed to answer, and it answers it well.

What the record does not independently settle is whether the target was what the targeting packet said it was. Battle-damage assessment can confirm an effect without confirming identity. It can show that a strike achieved its intended result without reopening whether the strike should have happened.

When an operator objects and the strike does not happen, there is no engagement to assess. If the operator was right, the intervention has erased its own evidence, leaving only a disagreement the record cannot settle. The institution still has to decide whether that operator's judgment was any good.

Park's file

Major Lena Park had ninety seconds to decide what the gathering was. The packet described vehicles arriving under a blackout, a fuel signature burning steadily, movement the system read as a logistics node assembling after dark. Park opened the underlying imagery and thought it looked like a burial.

She was not sure. The shapes were consistent with what the system said, and consistent with what she thought, and ninety seconds is not long enough to become certain of either. She entered the objection, chose a reason code from the menu the interface required, and deferred the target. The queue advanced. The next packet was already waiting.

Nothing about the objection was difficult. The system was built to accept it. There was no argument, no phone call, no senior officer asking her to justify herself. She clicked, the strike did not happen, and because it did not happen there would never be a record of what the gathering had actually been.

A few packets later the system did something she could not have done. It flagged a truck she had already dismissed and tied it, through a chain of movement and prior sightings, to a convoy the unit had been tracking for a week. She pulled the records. The link held. In ninety seconds she would never have found it, and the older way of working, the one with more humans and more time, had been slower without being obviously safer. She had watched that way miss things too. The system had earned her trust.

Months later, renewing a credential, she opened her own performance file.

The file reported her concordance with the system's recommendations, and the number was high. Most of the time she and the system agreed, which was unremarkable, because most targets are not close calls. What the file tracked was the small remainder where she had disagreed.

Some of those entries were settled against her. One was a target she had blocked that later intelligence confirmed had been valid. She had been wrong, the record said so, and the record was right.

Other entries came from objections that had been overridden. The strikes went ahead, and the assessment beside each one reported that the intended effect had been achieved. The building came down, the vehicle stopped, and nothing in those lines was false. But none of it addressed what Park had disputed. Her question was not whether the strike would work. It was whether the target was what the packet claimed. The assessment answered the question she had not asked and stayed silent on the one she had.

Then there was the gathering. She found it in the file as a deferred engagement, no adverse finding, no confirmation. The strike had not happened, so there was nothing to assess, so there was nothing beside the entry but the fact of her disagreement. The file could not tell her she had been right. It could not tell her she had been wrong. It could only show that she had stood in the way of something the system recommended, and that the something had not occurred, and that no one would ever know what it was.

She understood the argument for measuring reviewers. She had worked with one who flagged nearly every packet and held up the queue without catching more mistakes. An institution needed some way to distinguish careful judgment from indiscriminate objection. She had wanted the number to exist.

What she went looking for, in the credential system, was a way to reopen the calls that had never resolved. A place to enter what she had seen in the imagery, to have someone go back to the deferred targets and settle them one way or another.

With no process for revisiting those calls, her file could record that she had disagreed but not whether she had been right. It did not mark her wrong; it marked her discordant.

At the end of the next shift, the panel showed reviewer concordance within standard and no interventions for the period. The panel counted that as active human oversight.

Two ledgers, different clocks

The same objection can matter on two different timelines.

Park's objection entered the operational record immediately. That record could show that she had disagreed with the system, but not whether she was right.

The evidence that might settle the call moves more slowly, if it arrives at all. Operator feedback and combat imagery are documented inputs into how these systems keep developing. If later evidence established that a flagged pattern was civilian, that finding could inform a future version of the model. Whitworth, the NGA director, once called operator feedback part of "a perfect marriage" with the machine.

But the two clocks do not reconcile with each other. A confirmed pattern can improve the model going forward. It does not travel back to Park's file and correct the entry that scored her disagreement before anyone knew she was right. The model is designed to be updated. The personnel record is not, unless someone builds the path that would carry the correction backward, and nothing in the file suggests anyone did. The institution can get smarter and leave the person who made it smarter marked down for the disagreement that taught it.

None of this requires anyone to act in bad faith. The score accurately records agreement, and the assessment accurately records effect. Neither captures the judgment Park was asked to make. No individual step in the chain has to be corrupt for the gap to do its work.

A score wants nothing

AI-assisted targeting depends on human judgment by design. Operator feedback and combat imagery are documented inputs into continued development. Battle-damage assessment measures whether an engagement achieved its effect. The NGA has stood up an accreditation pilot for the methodology and testing behind its AI models. What the public record does not show is a process that scores individual reviewers this way, that adjudicates whether a given operator's call was correct before it becomes training data, or that carries such a correction back to the person who made it. The concordance file in Park's story is invented. The pieces it is assembled from are not.

The conditional matters. If an institution treats agreement with a recommendation as a measure of reviewer accuracy and validates it mainly through strike outcomes, the score inherits the outcome record's blind spots. It can see when a reviewer blocked a target later confirmed as valid. It cannot settle an objection that prevented an engagement unless the institution creates a separate way to revisit it. Without that path, the metric rewards agreement and marks down the disagreements it cannot resolve.

Asimov needed the Machines to decide who stood in their way. The story's threat depended on somebody wanting the outcome and concealing that intention.

A concordance score wants nothing. It only has to be built on the record the institution already keeps, and that record, on its own, cannot tell a target that was destroyed from a target that was correctly identified.