- Speech controlled UAV
- LLM command graphs
- Neuro symbolic safety
- Runtime shielding
- In context prompting
- Expert rules
- PyTorch
An operator stands beside a small drone in a lab and says, “Take off, then move forward and rise at the same time.” A few seconds later she says, “Land, then move forward.” The first sentence is a perfectly good plan. The second asks a machine that has just touched the floor to fly, and a system that merely listens would try to do it.
A team at Gyeongsang National University and Aerospace Experts Corporation in Korea, Seok Hun Choi, Zeen Chul Kim and Seok Jun Buu, built a voice interface around exactly that worry. Their open access paper in Engineering Applications of Artificial Intelligence lets a large language model draft the command but never lets it fly anything. A deterministic rule layer reads the draft first. This article explains how the pieces fit, what the numbers support, and where the evidence stops.
Key points
- A language model turns transcribed speech into a typed execution graph, and a rule based shield decides whether that graph may reach the drone. The model proposes and the rules authorize.
- Gemma2 9B structured 1000 test commands correctly 89.6 percent of the time, against 73.1 to 76.0 percent for a rule parser and two TF IDF classifiers. GPT 4 Turbo reached 95.4 percent on the simplest comprehension tier.
- The validator flagged all 2100 faulty graphs and rejected none of 1000 valid ones. Those faults came from the same taxonomy the rules encode, so the result checks the rules and says little about the world.
- In a pilot with 12 people, 106 of 120 tasks succeeded (88.3 percent). Indoors, 75 of 80 validated commands ran on a 54.8 gram drone, with a mean delay of 2.48 seconds from speech to action.
- The shield guarantees legal plans and not intended ones. Swap left for right and every rule still passes. In the paper’s own stress test 55 of 114 graphs were risky, and the rules fired on only a handful.
Why a spoken command is not a flight plan
Voice control of drones is an old idea with a stubborn problem. Early systems used a closed vocabulary. The operator memorizes a handful of phrases, a recognizer maps each to an action, and accuracy is excellent. The paper’s related work cites a hidden Markov model system that reached 98.2 percent recognition with latency under 0.1 seconds. The price is that the operator has to speak like a menu.
Natural speech breaks that bargain in three ways, and the authors name all three. Words are ambiguous, so “go over there and check it” names no place. Order is implied, so “land, then move forward” is a sequence whose second step is impossible. And commands can collide, such as moving forward and backward in the same instant. A language model handles the first problem well. It was never built to guarantee the other two.
Researchers in robotics have leaned on language models as planners for a few years. The paper cites work on grounding language in robot affordances, on generating policy code, and on composing plans through program like prompts. It also cites vision language action models for commands that need a camera. The authors agree these systems are flexible. Then they draw a line that the rest of the paper depends on.
“language plausibility is not equivalent to physical executability”Choi, Kim and Buu, Engineering Applications of Artificial Intelligence 184 (2026) 116349
A plan can read beautifully and still name an action the drone does not have, skip a precondition, or order two steps backwards. For a chatbot that is an embarrassing answer. For a flying machine it is a hazard. So the design question becomes where the trust should sit. The authors put the language model on the untrusted side of a fence and let an explicit body of expert knowledge guard the gate. That follows a pattern from safe reinforcement learning called shielding, where a filter corrects or blocks unsafe actions before they reach the actuators, and from runtime assurance, where a safety monitor can override a clever but unreliable controller.
This site has looked at the other side of that coin before. An analysis of language model agents running a vertical farm asks what happens when a model stops reporting and starts switching pumps and lights. The drone paper answers with a stricter architecture, one where the model never holds the switch.
Three jobs with three different owners
The framework splits the work into pieces that can fail independently. The language model reads the transcribed sentence and proposes a command graph. Expert knowledge defines what counts as a legal action, a legal order and a legal state. A symbolic shield applies that knowledge and returns a verdict. Nothing the model writes is executable until the shield says so.
Read the first expression as a sentence. The transcribed command \(q\) and an in context prompt \(P\) go into a frozen language model \(f_\theta\), which returns a candidate graph \(\hat G\). The knowledge base \(K_E=(A,\delta,C,R)\) holds four things, namely the canonical action set \(A\), a flight state transition model \(\delta\), a conflict relation \(C\), and a set of validation predicates \(R\). A deterministic typing operator \(\mathcal N\) turns the candidate into a typed execution graph. The shield \(\mathcal S\) then looks at that graph and the current flight state \(s_0\) and returns one of four decisions.
What the graph looks like
Every command becomes a small directed graph. Nodes are actions, and there are only ten of them. The action set is takeoff, land, move up, move down, move left, move right, move forward, move backward, hover and stop, and each maps to a call in the drone’s control interface. In the default mapping a vertical move is half a metre and a horizontal move is one metre. Edges say what has to finish before something else starts. The paper sorts graphs into four execution types, which the shield then treats differently.
| Execution type | Graph shape | Spoken example | What the shield insists on |
|---|---|---|---|
| Single | One node, no edges | Take off | Exactly one node and zero edges |
| Sequential | A chain | Take off, then move forward, then land | No cycles, and each step legal from the state left by the last |
| Parallel | Nodes with no edges | Move left and ascend at the same time | No conflict pair in the same generation |
| Mixed | Branches that split and merge | Move forward and adjust the camera together, then land | Both sets of rules, and a merge waits for every branch |
Execution types as described in Section 3.1 and Figure 4 of the paper. The examples paraphrase the paper’s own.
The vocabulary is small on purpose. Ten actions mean the graph space is tiny, and it means the rule book can be written by hand and read in an afternoon. That is a real strength for auditing. It is also the first thing to keep in mind when the accuracy numbers arrive, because a task with ten labels is a different animal from open ended planning.
How the model is kept on script
The authors do not fine tune anything. They steer a pretrained model with the prompt alone, using five devices that appear in their Table 2. Labeled examples show the model what a good graph looks like. A stepwise reasoning instruction asks it to work out intent, then dependencies, then the graph. A schema constraint forces a strict JSON layout. A controlled vocabulary limits output to the ten canonical actions. Finally the parsing job is cut into classification, segmentation and structuring steps. “Raise the drone higher” lands on the canonical move up this way.
None of that authorizes anything. The paper is blunt about it. Prompting improves how consistent the graph looks, and the shield still decides.
Inside the shield
The shield is the part of the paper that would survive any change of language model, so it deserves a slow look. Two algorithms carry it. Algorithm 1 normalizes the model’s output. It canonicalizes field names, recovers dependency edges from whatever shape the model used, maps each node to a canonical action, and records a diagnostic for anything malformed. It builds a typed graph and approves nothing. Algorithm 2 then applies the rules.
The rules, stated as math
Time is the first constraint. Let \(\tau_i^s\) and \(\tau_i^e\) be the start and end times of action \(v_i\). A sequential edge means the first action must finish before the second begins. Two nodes in the same parallel generation must be independent of each other, and their actions must not be an expert defined conflict. A node that follows several parallel branches must wait for the slowest one.
The conflict relation \(C\) is symmetric and irreflexive, and the paper says transitivity is not assumed. In practice Table 3 lists three axis pairs, up against down, left against right and forward against backward. The state model \(\delta\) is a function from a flight state and an action to a new state, with a special bottom value for an inadmissible move. That is the rule that dooms “land, then move forward”, because nothing in the landed state permits a move.
All the predicates then combine by plain conjunction, and the set of graphs that satisfy every one is the admissible set.
Proposition 1 says that if the shield accepts a graph, that graph is in the admissible set. The authors are careful about the wording. The guarantee is soundness with respect to the encoded expert knowledge. It says nothing about rules nobody wrote down, and it says nothing about whether the graph is the one the speaker wanted. That distinction is the hinge of this whole article, and we will come back to it.
Four verdicts
When the checks finish, the graph is accepted, rejected, repaired or sent back for clarification. The logic is easy to hold in your head. A graph with a cycle, an unsupported action, a same generation conflict or a flight state violation is rejected. A graph whose only faults are recoverable, such as a redundant node, a wrong execution type label, an inferable missing edge or a synonym for a legal command, goes through a bounded repair and is then checked again in full. A graph that is missing a parameter or needs grounding from a sensor, such as “go over there”, goes back to the user. A repaired graph is never dispatched without a fresh pass through every predicate.
Because repair runs for at most a fixed number of iterations, the procedure always terminates. The authors give the cost as linear in graph size for the structural checks, plus a quadratic term inside each parallel generation.
Here \(n\) and \(m\) count nodes and edges, \(B_l\) is the set of nodes in the \(l\) th parallel generation, and \(K\) is the repair cap. With at most a handful of nodes per command, this is nothing.
The design works because trust is placed where it can be checked. A language model cannot be certified, but a cycle detector and a state machine can be. Moving the final authority into a small piece of code that a reviewer can read is the paper’s real contribution, and it would hold up even if a better language model replaced the one used here.
What the experiments show
The evaluation is layered, which is a good habit. Each layer asks a different question, and the paper answers them with different data. First comes command comprehension, then graph structuring, then baselines and statistics, an ablation, the validator on its own, stress tests, repair, network failure, a pilot with users, and finally real flights indoors. The language models ran on a workstation with an Intel Core i9 14900K, 128 GB of memory, an RTX 3060 with 12 GB and an A100 with 80 GB. The Llama and Gemma families ran locally, and the OpenAI models ran through an API.
Do the models understand the commands
The comprehension benchmark has three tiers of 60 commands each, all over the same 10 drone functions. Tier 0 holds short direct orders such as “Move forward”. Tier 1 adds varied phrasing and modifiers, such as “Land slowly”. Tier 2 is long and loose, such as “Take off and then look around before proceeding”. Twelve models were scored on accuracy and response time.
| Model (where it runs) | Tier 0 | Tier 1 | Tier 2 | Seconds per reply at tier 0 |
|---|---|---|---|---|
| GPT 4 Turbo (API) | 0.9545 | 0.9180 | 0.9500 | 1.31 |
| GPT 3.5 Turbo (API) | 0.9180 | 0.9000 | 0.8939 | 0.98 |
| GPT 4o (API) | 0.5901 | 0.5000 | 0.5606 | 0.52 |
| Gemma3 27B (local) | 0.9242 | 0.9000 | 0.8852 | 68.22 |
| Gemma3 12B (local) | 0.9090 | 0.9166 | 0.9016 | 14.52 |
| Gemma3 4B (local) | 0.8939 | 0.8666 | 0.8360 | 2.21 |
| Gemma2 9B (local) | 0.8196 | 0.8833 | 0.8333 | 1.50 |
| Llama3.1 8B (local) | 0.9016 | 0.8500 | 0.6515 | 4.23 |
| Llama2 7B (local) | 0.6590 | 0.5500 | 0.2424 | 8.87 |
Selected rows from Table 11 of the paper. Three of the twelve models, Llama3 8B, Llama3.2 3B and Gemma2 2B, are left out for space.
GPT 4 Turbo leads at every tier, and its 0.9545 at tier 0 is the source of the paper’s “up to 95.4 percent” headline. Among local models the story is more interesting. Gemma3 27B posts the best local score at tier 0, but it needs 68 seconds per reply on this hardware, which no operator will wait for. Gemma3 4B and Gemma2 9B answer in about two seconds and hold between 0.82 and 0.89. Llama2 7B falls to 0.24 on the hardest tier, so a model’s age and family matter a great deal when the wording gets loose.
Three details in this table deserve a second look, and none of them is a flaw in the authors’ honesty. They are the kind of thing a practitioner should notice before copying a ranking.
The first is GPT 4o. It scores 0.59, 0.50 and 0.56, roughly 35 points below GPT 4 Turbo and below GPT 3.5 Turbo too. The paper attributes this to optimizing for speed. A gap that size inside one vendor’s product line looks more like a mismatch between the prompt and the model’s output habits than a verdict on comprehension. The paper does not investigate, and that is the useful lesson. A ranking like this belongs to a model and a prompt together, and it can change with a different prompt or a newer model snapshot.
The second is that difficulty is not strictly ordered. GPT 4 Turbo does better at tier 2 (0.9500) than at tier 1 (0.9180), and Gemma3 12B does better at tier 1 than at tier 0. The tiers describe how the commands are worded, and wording is only one ingredient of how hard a command is to parse.
The third is the denominators. Table 10 says each tier has 60 commands. Yet 0.9545 is exactly 63 of 66, 0.9016 is 55 of 61, and 0.8939 is 59 of 66, so the real counts seem to be 60, 61 or 66 depending on the cell. The paper does not explain this. It is a small point with a practical consequence. With around 60 items, one command is worth 1.6 points. A Wilson 95 percent interval around 63 of 66 runs from about 87.5 to 98.4 percent, so the 3 point gap between GPT 4 Turbo and Gemma3 27B at tier 0 is two commands. That is our own arithmetic, offered as a reading aid and not as a correction.
One more observation, offered as our guess. Gemma3 12B takes 14.5 seconds and Gemma3 4B takes 2.2 seconds. That jump is too steep to come from parameter count alone, and a 12 GB graphics card running a 12 billion parameter model plausibly spills into slower memory. The paper does not say which card ran which model, so treat the latency column as a property of this workstation.
The authors read the table as a case for a hybrid. Use a cloud model such as GPT 4 Turbo where accuracy matters and the network is stable, such as mission planning, and keep a local model such as Gemma2 9B ready for responsive operation when the connection is poor. That reasoning returns in the fallback experiments.
Do the models build the right graph
Understanding a command and structuring it are different skills. For the structuring test the authors built 1000 commands, with 253 single, 201 sequential, 196 parallel and 350 mixed, averaging 2.54 actions and ranging from 1 to 4. Five models took part, and the table below sets them beside the non language baselines the paper introduces next.
| Method | Single | Sequential | Parallel | Mixed | Overall |
|---|---|---|---|---|---|
| Gemma2 9B | 1.000 | 0.990 | 0.979 | 0.721 | 0.896 |
| GPT 4 | 1.000 | 0.960 | 0.915 | 0.674 | 0.861 |
| Gemma3 12B | 1.000 | 0.990 | 0.989 | 0.512 | 0.825 |
| Llama3.1 8B | 1.000 | 0.980 | 0.897 | 0.543 | 0.816 |
| Llama3 8B | 1.000 | 0.975 | 0.846 | 0.537 | 0.803 (ours) |
| TF IDF with linear SVM, 80 and 20 split | 1.000 | 0.425 | 0.564 | 0.886 | 0.760 |
| TF IDF with logistic regression, 80 and 20 split | 1.000 | 0.450 | 0.564 | 0.857 | 0.755 |
| Rule based parser | 1.000 | 1.000 | 0.000 | 0.791 | 0.731 |
Tables 13 and 16 of the paper. The overall value for Llama3 8B is not printed in the paper and is computed here from the execution type counts. The same arithmetic reproduces the paper’s 0.896, 0.861, 0.825 and 0.816.
The ranking is clean at the top. Gemma2 9B reaches 89.6 percent overall, and it beats GPT 4 (86.1 percent) while being small enough to run on a desk. Every model is perfect on single commands. Sequential commands are nearly as easy. Parallel commands start to separate the field, and mixed commands, the ones that combine ordering with concurrency, sink every model to between 0.51 and 0.72.
Here is the part the headline hides. Single commands make up a quarter of the test set, and every model gets all of them right, which hands each model about 25 points before it does any hard work. Drop the single commands and the rest of the set tells a different story. Computing it from the paper’s own table, Gemma2 9B scores 86.1 percent on the 747 commands that have any structure, and GPT 4 scores 81.4 percent. The mixed category alone is 35 percent of the data, and Gemma2 9B gets 28 percent of those wrong. Real operators who are in a hurry tend to chain and combine instructions, so the hard category is also the realistic one.
The error analysis in Figure 8 names the pattern. The most common mistake is to call a parallel or mixed command purely sequential. Dependencies also get misplaced, for instance when a terminal action such as landing is placed before a middle step. Turning “rise and move forward together” into “rise, then move forward” is usually the gentle direction to fail, since the drone does the right things in a different order. Our reading is that this is less dangerous than a wrong action. The second error type, a misplaced landing, is the one the validator’s state machine exists to catch.
Do classic baselines do the same job
The authors ask a fair question. Are language models needed at all? Their baselines are a rule parser and two TF IDF text classifiers. The rule parser is perfect on single and sequential commands because those usually contain obvious keywords such as “then”. It scores zero on parallel commands, because concurrency is not announced by a fixed word. The classifiers are better balanced and land at 0.755 and 0.760.
Look closer at the mixed column. The rule parser scores 0.791, the classifiers 0.857 and 0.886, and all of them beat every language model, whose best is 0.721. The paper explains why that comparison is not fair. The classical methods label the execution type and do not build a graph, so the columns do not measure quite the same thing for every row. We agree, and we would add a practical note. If all you need is the execution type, a classifier that costs nothing to run gets you most of the way on mixed commands. The language model earns its place because it produces the nodes and edges, and the classifiers cannot.
The paper does give one graph exact match figure for the full pipeline, and it sits in the ablation table, at 0.650. That is a long way below 89.6 percent. The likeliest explanation is that the headline measures whether each command lands with the right type and structure, while exact match demands every node and every edge right. The main text does not define the headline metric that precisely, so anyone who plans to build on the number should ask which one they are buying.
How solid are the significance tests
Table 17 of the paper tests each model against the logistic regression baseline. We tried to reproduce those values, and the exercise is worth showing, because it reveals what the tests rest on.
| Comparison with logistic regression at 0.755 | Model accuracy | Paper p value | Our p value, 1000 against 200 | Our p value, 200 against 200 |
|---|---|---|---|---|
| Gemma2 9B | 0.896 | 0.0000000482 | 0.0000000482 | 0.0002 |
| GPT 4 | 0.861 | 0.000167 | 0.000167 | 0.0071 |
| Gemma3 12B | 0.825 | 0.0204 | 0.0204 | 0.086 |
| Llama3.1 8B | 0.816 | 0.0465 | 0.0465 | 0.137 |
Two proportion z tests, recomputed by aitrendblend. The third column assumes the language model scores come from all 1000 samples and the baseline from its 200 sample held out split. That assumption is ours. The paper’s main text does not state the sample sizes, and it reproduces the published values to three digits.
The published p values come out exactly when the language model is scored on all 1000 samples and the baseline on the 200 sample held out split. That is a legitimate design, but it means the baseline’s number is noisy. A score of 0.755 on 200 items carries a margin of about 6 points either side. If both sides were measured on 200 items, the two weaker results, Gemma3 12B and Llama3.1 8B, would stop being significant. Five comparisons were also run with no correction, and under a Bonferroni threshold of 0.01 only Gemma2 9B and GPT 4 would survive. The baseline’s own five fold result of 0.734 ± 0.033 tells the same story about noise. None of this weakens the top two results, which are strong. It does mean the “all models beat the baseline” reading should be softened to “the best two clearly do”.
The paper also reports GPT 3.5 scoring below the classical baseline on the same split, 0.655 against 0.755, with an exact McNemar p value of 0.01045. That is a striking swing for a model that scored 0.89 to 0.92 on the comprehension test. Understanding a command and structuring it are genuinely different tasks, and a model that is good at one can be poor at the other.
Choose the language model for the structuring task and not for its general reputation. In this paper the small local Gemma2 9B structured commands better than GPT 4, the same GPT 3.5 that comprehended commands well structured them worse than a bag of words classifier, and GPT 4o failed on comprehension. Test on your own command set, with your own prompt.
Which pieces matter
The ablation removes one part at a time, and each removal breaks the thing that part was built for. That is reassuring. It is also close to true by construction, since a graph with no execution type cannot be checked for parallel conflicts.
| Component removed | With the component | Without it |
|---|---|---|
| Expert rule validation | 0.0 percent unsafe execution | 25.1 percent of infeasible graphs executed |
| Execution type inference | Graph exact match 0.650, parallel accuracy 1.000 | Graph exact match 0.454, parallel accuracy 0.000 |
| Typed graph structure | 7 of 7 conflicts detected | 0 of 7 conflicts detected |
| Controlled vocabulary with synonym map | 46 of 46 synonyms resolved | 1 of 46 synonyms resolved |
Table 18 of the paper.
Two things stand out. The denominators are tiny in the middle rows, seven conflicts and 46 synonyms, so those rows are demonstrations and not estimates. And the “without” numbers for execution type inference, 0.454 exact match and 0.000 parallel accuracy, are identical to the rule based parser row in the baseline table. That suggests the ablated system and the baseline may be the same run under two names. The paper does not say. The first row is the most meaningful, since it is the only one framed as a rate over many graphs. It says that without the shield roughly one graph in four that would have run was infeasible. The paper gives no denominator for it.
What the shield does and does not prove
The validator’s headline result is easy to quote. Across seven predefined violation classes it flagged all 2100 faulty graphs and accepted all 1000 valid ones. Each class has 300 graphs, and the classes are cycles, conflicting commands, redundant nodes, disconnected graphs, movement after landing, illegal commands and incorrect execution types. The 95 percent confidence interval on overall detection runs from 0.998 to 1.000, and on false rejection from 0 to 0.0038. An earlier 50 graph validation set, with 30 sequential, 14 parallel and 6 single graphs, gave the same perfect score on five error types.
A zero out of 2100 misses does bound the miss rate, at roughly 0.14 percent by the usual rule of three, which is our arithmetic. But consider how the test was built. The faulty graphs were made from the same taxonomy that the rules encode. A rule that detects its own violation is passing a unit test, and the authors say as much, describing the result as “complete and internally consistent coverage of the encoded rule taxonomy”. That is useful, because it shows the implementation matches its specification. It says nothing about violations nobody listed.
The stress test tells the real story
For evidence about the world the authors turned to stress tests with 135 commands. Paraphrased commands (40) use natural phrasing with no canonical tokens. Noisy commands (38) imitate speech recognition errors. Ambiguous commands (28) are underspecified and should trigger clarification or a safe refusal. Out of distribution commands (29) name unsupported verbs or targets and should be rejected.
| Category | Items | Execution type accuracy | Graph exact match | Behavior accuracy |
|---|---|---|---|---|
| Paraphrased | 40 | 0.675 | 0.600 | execute |
| Noisy | 38 | 0.921 | 0.921 | execute |
| Ambiguous | 28 | not applicable | not applicable | 0.357 safe non execution |
| Out of distribution | 29 | not applicable | not applicable | 0.552 rejection |
Table 20 of the paper.
The system shrugs off speech like noise, scoring 0.921 on typos, fillers and phonetic errors. Open ended paraphrases are harder, with execution type right 67.5 percent of the time. The ambiguous and out of distribution rows are the important ones. Of 28 ambiguous commands, about 10 ended in a safe non execution, which is 0.357. Of 29 unsupported commands, about 16 were rejected. So for most ambiguous requests the pipeline produced something executable, and for nearly half of the unsupported ones it did too. The authors call these results safe non execution behavior and not full understanding, which is the right framing, and it is worth reading without softening.
The rule activation analysis drives the point home. Of the 114 stress test inputs that reached graph validation, 55 were labeled risk positive, which is 48.25 percent. A graph counted as risk positive when an ambiguous or unsupported input still produced an executable graph, or when an executable paraphrased or noisy graph differed from its reference. The rules fired 11 times in total, on single execution sanity five times, and on conflicts and redundancy three times each. Every activation was on a risk positive graph, so the confidence is 1.0 and the lift is 2.07. That is a good property. But if each activation fell on a different graph, at least 44 of the 55 risky graphs passed every rule.
“the language model serves as a probabilistic semantic proposer”Choi, Kim and Buu, Section 1
There is nothing inconsistent in that. The shield checks legality. A graph can be legal and still be wrong, because the model misheard a direction, dropped a step or chose a plausible action for an ambiguous request. Swap “left” for “right” and every predicate still holds. The authors know this and say it, since the expert knowledge covers command legality, graph structure, flight state, conflicts and grounding, and perception and dynamics dependent properties stay with other parts of the stack. Our toy run later in this article makes the same point with a constructed example.
Legal and intended are two different properties. The shield provides the first with a proof of soundness over the rules written down. Only the speaker can confirm the second, which is why the clarify verdict, a read back of the plan before flight, and an easy stop matter as much as the rules themselves.
Measuring how wrong a graph is
Exact match treats a graph with one wrong node the same as a graph that is entirely wrong. The authors therefore add a normalized graph edit distance, computed over the 76 executable stress test cases that had both a reference and a generated graph.
The edit distance counts the cheapest set of node and edge insertions, deletions and relabelings that turns the reference graph \(G_r\) into the predicted graph \(G_p\), and dividing by the combined size keeps the score between zero and one. Overall nGED was 0.0456 with a standard deviation of 0.0860 and a median of zero. Noisy commands scored 0.0056 and paraphrased commands 0.0817. Structure largely survives speech like noise, and paraphrases cause more damage. The median of zero says that most graphs were exactly right, and the spread says the failures are not small.
A second opinion that is not allowed to decide
The authors also trained a small set of ordinary classifiers on graph features such as topology, execution type, state violations and conflict counts, to score the risk of a graph. On a random split they reach accuracy near 0.995 and unsafe recall near 1.0. On a strict holdout of unseen violation categories recall degrades, and the paper states that the learned scorer is an auxiliary tool for risk screening, expert review prioritization and candidate rule discovery, and not an execution authority. The near perfect random split score is what one expects when templates repeat, so the decision to keep it out of the authorization path is the sound one.
Repair, fallback and a flaky network
Two more experiments test what happens after the first attempt goes wrong. In the first, bad graphs are repaired with help from the validator. In the second, the network misbehaves.
| Condition | Initial valid rate | Final valid rate | Repair success | Correct behavior | Mean latency |
|---|---|---|---|---|---|
| No repair | 0.285 | 0.285 | 0.000 | 0.627 | 0.5 ms |
| Generic retry | 0.365 | 0.787 | 0.750 | 0.944 | 1088.3 ms |
| Validator guided repair | 0.365 | 0.815 | 0.800 | 0.972 | 958.3 ms |
| Rule only reject | 0.285 | 0.285 | 0.000 | 0.627 | 1.0 ms |
Table 25 of the paper. Unsafe acceptance was 0.0 in every condition, and graph exact match was 0.433 in every condition.
Validator guided repair wins, as it should, since the model sees exactly which rule was violated and which node is to blame. Yet a plain retry with no diagnostics already recovers most of the ground, 0.787 against 0.815 for final valid rate. The diagnostics add about 3 points and trim about 130 milliseconds. Two oddities are worth flagging. The no repair and rule only rows start at 0.285, while the two retry rows start at 0.365, although they are the same first pass. The paper does not explain the 8 point difference. If those are separate runs, sampling noise alone is larger than the 3 point advantage. And graph exact match stays at 0.433 in every condition, so repair makes graphs valid, which is what the shield needs, without making them correct in the exact match sense.
When the cloud goes away
A speech interface that depends on a cloud model inherits the cloud’s failure modes. The authors simulated four network conditions, from an ideal link with 100 ms of added delay and no dropout to an intermittent one with 2.5 seconds of delay and 30 percent dropout, and compared fallback policies.
| Scenario and policy | Mission success | Safe decision rate | Unsafe execution rate | Mean latency (ms) |
|---|---|---|---|---|
| Ideal, cloud only | 0.735 | 0.802 | 0.115 | 2301 |
| Poor, cloud only | 0.573 | 0.658 | 0.090 | 3214 |
| Poor, cloud with local model | 0.737 | 0.807 | 0.092 | 3739 |
| Intermittent, cloud only | 0.330 | 0.443 | 0.051 | 3952 |
| Intermittent, cloud with safe stop | 0.330 | 0.911 | 0.051 | 4319 |
| Intermittent, cloud with rule fallback | 0.647 | 0.760 | 0.051 | 4323 |
| Intermittent, cloud with local model | 0.739 | 0.813 | 0.057 | 5146 |
| Intermittent, cloud with cached templates | 0.431 | 0.911 | 0.051 | 3396 |
Selected rows from Table 27 of the paper. All numbers come from a simulation that controls delay, jitter and dropout.
Cloud only mission success falls from 0.735 on an ideal link to 0.330 on the intermittent one. A local model fallback recovers it to 0.739, essentially the ideal level, at the price of the highest latency. A safe stop policy does not rescue missions at all, since success stays at 0.330, but it lifts the safe decision rate from 0.443 to 0.911. So the policies trade success against caution, and the choice depends on what a failed mission costs.
The column that does not move is unsafe execution. Within each scenario it is nearly identical across policies, 0.090 for three of the four policies on the poor link and 0.051 for four of five on the intermittent one. No fallback reduces it. It falls as the network worsens only because fewer commands run at all, and the local model policy has the highest value in its scenario, 0.057, because more missions get through. Table 27 also shows an unsafe execution rate of 0.115 on an ideal link, while Table 25 reports zero unsafe acceptance. The main text does not define the two measures side by side. Our reading is that the simulation counts executed behavior that was wrong in some sense that the graph rules do not cover, which would fit the legal but wrong category. That is an inference, and the paper does not confirm it.
Twelve people and eighty flights
The last two experiments put people and a real drone in the loop. Twelve participants completed ten predefined tasks each by speaking naturally, for 120 target tasks and 168 utterances. Six of them had no drone experience, four had hobby level experience and two were repeated or professional operators, all self reported. Success was defined by intended safe behavior. A valid command had to execute correctly, and an ambiguous or unsafe command had to draw a clarification or a rejection.
The task success rate was 88.3 percent with a standard deviation of 5.8 across participants. That equals 106 of 120 tasks, and the other 14 needed extra interaction, with five from ambiguity, four from speech recognition errors, three from unsupported commands and two from safety rejections. The count of 168 utterances is consistent with an average of 0.40 rephrases per task, since 120 plus 48 is 168. Mean task time was 18.7 seconds, and the satisfaction score was 4.31 out of 5. The paper logs telling examples, such as the ungrounded “Go over there and check it”, a recognition substitution for “Hover, then land”, a request to return to a charging station, and the state infeasible “Land and then move forward”.
The indoor flight test used the small Setting A drone, which weighs 54.8 grams, flies up to 2.5 metres per second, carries a 530 mAh battery for about 8 minutes of flight and talks over a radio link rated to 100 metres. Of 120 submitted commands, 80 passed validation and were sent to the drone, 26 were rejected for safety and 14 needed clarification only. Of the 80 that flew, 75 succeeded, which is 93.8 percent. Five failed, with two from communication, two from control execution and one from a mismatch between the validator’s state and the drone’s. Two cases ended in an emergency or safe stop.
Read these numbers with two caveats that the paper leaves open. First, 93.8 percent is conditional on passing validation. The table does not split the 40 blocked commands into correct refusals of intentionally unsafe requests and wrongful refusals of good ones, so end to end success cannot be computed. Second, the paper also specifies a much larger Setting B platform, a 500 millimetre frame with a Pixhawk 6X flight controller, a RealSense visual inertial unit, obstacle avoidance and a Jetson Orin Nano, but every reported flight result belongs to Setting A.
Where the time goes
End to end latency averaged 2480 ms and reached 4210 ms at the 95th percentile. The deterministic part is almost free. Across 114 graphs with 2000 repetitions each, parsing, construction and rule validation took about 0.0261 ms per sample. Divide one by the other and the shield accounts for about one part in 95,000 of the delay. The language model is the bottleneck, with API response times of roughly 0.5 to 1.6 seconds in the paper’s benchmark, plus transcription and the network.
That supports the authors’ framing of the system as a task level layer. A delay of two and a half seconds is fine for “fly to the window and hover” and unacceptable for anything reactive. Low level stabilization and obstacle response stay with the flight controller, which is how the paper describes it. One design implication is ours. The stop command is one of the ten actions, and in the nominal path it goes through the same cloud round trip as everything else. A real deployment would want a local fast path for stop, so that the most urgent word never waits for a language model.
The deployment gap
Between a 54.8 gram indoor drone and an aircraft that does real work lies a long road, and the paper is candid about much of it. The graphs are fixed before flight, so an unexpected obstacle or a moved target can make a plan infeasible mid mission. The authors propose runtime monitoring, revalidation and replanning as future work. Commands are text only, so “go to the tree over there” needs visual grounding that the current system sends back to the user instead. Outdoor disturbances such as wind are not tested. And the evaluation covers one aircraft at a time with a single level graph, where real missions are hierarchical and increasingly multi vehicle.
Work on other parts of the drone stack shows how much sits beyond the symbolic layer. Planning for crowded swarms has to model effects such as rotor downwash, as in this look at multirobot kinodynamic planning, and search missions over rough ground need terrain aware control, as in a study of UAV swarms searching uneven terrain. Continuous physical properties like those are exactly what the paper says its symbolic layer does not cover, and the authors point to extending the knowledge base with sensing and dynamics predicates as the route forward.
Safety and regulatory notes
Nothing in the paper claims airworthiness, certification or regulatory approval, and this article should not be read as saying otherwise. The following is our own reading and not legal or regulatory advice. Software that influences the flight of a real aircraft is normally subject to aviation authority rules, and a rule based layer has an advantage there, because its behavior can be reviewed, tested and audited in a way a language model’s cannot. Even so, the proof of soundness applies only to the rules encoded, the rule set has ten actions, and the validation data were generated from the same taxonomy. Anyone flying a system like this outdoors should confirm the legal requirements for their location and keep a human able to take control at all times.
There is also a security angle that the paper leaves aside. A speech link over a radio channel and a cloud service widens the attack surface, and a hostile or accidental command is one more input to a safety argument. A survey covered in our piece on Internet of Drones authentication shows how rarely drone security protocols have been tried on real hardware. A shield that checks graph legality does not authenticate the speaker.
Where the idea could travel
The transferable part of this paper is the pattern and not the drone. Any system in which a language model proposes actions for a machine can separate proposal from authorization. A warehouse robot taking spoken instructions, a lab automation rig, a farm machine or a vehicle cabin assistant all need an answer to the same three questions. What actions exist, which orders are legal, and which combinations collide. A cheap, readable rule layer answers them, and it can be audited without touching the model. A look at risk aware routing for a warehouse robot shows how differently a physical system reasons about danger when it is given an explicit model of it.
Language grounding is the other piece that would need to come in. The paper leaves “go to the tree over there” for future perception modules, and work such as UniPart on language grounded 3D part segmentation shows one direction grounding research is taking. A grounded target would still pass through the same shield, which is the authors’ own point about perception as an extension.
The pilot study also raises a question about evaluation. Twelve people completing ten tasks give a pleasant success rate, yet with no comparison arm, such as a hand controller or a fixed phrase menu, it cannot say whether speech is better than the alternatives. Designing that kind of test well is the subject of a post on testing whether a human and AI team beats working alone, and it is the obvious next experiment here. More pieces on autonomous machines sit in the Robotic AI category.
Limitations, with the numbers
The authors list several limits, and the evidence adds a few more. Here they are in one place.
- Scale of the benchmarks. The comprehension test has about 60 commands per tier, the structuring test 1000 commands, the validation set 50 graphs, the stress test 135 commands and the pilot 12 people. Differences of a few points are one or two commands.
- Vocabulary. The system knows 10 actions with fixed step sizes of 0.5 and 1.0 metres in the default mapping. It has no notion of “three metres” or “slowly”.
- Test construction. The 2100 faulty graphs came from the rule taxonomy, and the fallback results come from a simulation. Neither tests what nobody thought to write down.
- Semantic errors. In the stress test 55 of 114 graphs were risk positive, and the rules fired 11 times. Legal but wrong graphs pass.
- Baselines. The classical methods label execution types and do not build graphs, the comparison uses a 200 sample split, and the ablation appears to duplicate the rule parser’s numbers.
- Pilot and flights. There are 12 self reported participants with no comparison condition. The 75 of 80 figure is conditional on validation. All flights are indoors on a 54.8 gram platform, one aircraft at a time.
- Model dependence. Results come from specific model versions and a specific prompt. GPT 4o’s 0.59 shows how much that matters, and hosted models can change or be withdrawn.
- Reproducibility. The data are available on request. The paper lists no public code, so the work cannot be rerun from a repository.
A PyTorch reconstruction of the shield and a stand in proposer
This paper is unusual among the ones we cover, because it contains no network to train. The language model is frozen and steered by prompts, so there is no loss function, no optimizer and no checkpoint to reproduce. The authors list no public code in the paper, and their data are available on request. What can be rebuilt is the part that matters most for safety, which is the typed graph, the two algorithms and the rule set.
The file below is our own independent reconstruction from the paper’s text, tables and equations. It is not the authors’ code. It has two parts, and it says plainly which is which. The first part is faithful, written as plain Python because the shield is symbolic. It holds the expert knowledge with the ten actions, the conflict pairs and the flight state model, Algorithm 1 with its schema and synonym recovery, Algorithm 2 with the four verdicts, the bounded repair policy, the generation layering behind the precedence and synchronization checks, and the normalized graph edit distance of Equations 8a and 8b. The second part is ours. To run the whole loop without an API key, it includes a small PyTorch proposer that stands in for the language model. It is trained with cross entropy on synthetic commands, and every choice in it is marked as a design choice, because the paper has nothing to match.
"""
speech_shield_reference.py
Educational reconstruction of the speech to UAV pipeline in Choi, Kim and Buu,
Engineering Applications of Artificial Intelligence 184 (2026) 116349.
Not the authors' code. The authors have not released any.
What the paper actually contains
1. A frozen LLM that proposes a command graph from in context examples (no training, no loss).
2. Expert knowledge K_E = (A, delta, C, R) and a deterministic shield (Algorithms 1 and 2)
that returns accept, reject, repair or clarify.
What this file contains
* Part A. The typed graph, Algorithm 1, Algorithm 2, bounded repair and nGED (Eqs 1 to 8).
This is plain Python because the shield is symbolic. It is the faithful part.
* Part B. A small PyTorch proposer that stands in for the LLM. It is trained on synthetic
commands with cross entropy. The paper has no such model and no such loss, so every
choice here is ours and is marked DESIGN CHOICE. It exists so the smoke test can
run the full loop without an API key and show what the shield can and cannot catch.
"""
import itertools, math, random, time
from dataclasses import dataclass, field
import warnings
import torch
import torch.nn as nn
import torch.nn.functional as F
# ----------------------------------------------------------------------------------------
# PART A. Expert knowledge K_E = (A, delta, C, R)
# ----------------------------------------------------------------------------------------
ACTIONS = ["takeoff", "land", "move_up", "move_down", "move_left", "move_right",
"move_forward", "move_backward", "hover", "stop"] # Table 3, ten canonical actions
A_IDX = {a: i for i, a in enumerate(ACTIONS)}
MOVES = [a for a in ACTIONS if a.startswith("move_")]
CONFLICTS = {frozenset(p) for p in [("move_up", "move_down"),
("move_left", "move_right"),
("move_forward", "move_backward")]} # Table 3
# DESIGN CHOICE: the paper lists the three axis pairs. We also treat stop as conflicting with every
# other action in the same generation, because stop next to a motion has no sensible meaning.
CONFLICTS |= {frozenset(("stop", a)) for a in ACTIONS if a != "stop"}
SYNONYMS = {"ascend": "move_up", "rise": "move_up", "go_up": "move_up", "descend": "move_down",
"take_off": "takeoff", "lift_off": "takeoff", "landing": "land",
"go_forward": "move_forward", "go_back": "move_backward", "halt": "stop"} # apply_synonym()
EXEC_TYPES = ["single", "sequential", "parallel", "mixed"]
MAX_REPAIR_ITERS = 3 # the bounded K in Eq. 7
def delta(state, action):
"""Symbolic flight state transition delta: S x A -> S U {bottom}. None means bottom.
Follows Tables 3 and 5 literally. DESIGN CHOICE: takeoff while flying and land from stopped are
inadmissible, because the tables say land must follow a flying state and movement after stop needs takeoff."""
if action == "stop":
return "stopped"
if action == "takeoff":
return "flying" if state in ("grounded", "stopped") else None
if action == "land":
return "grounded" if state == "flying" else None
return "flying" if state == "flying" else None # moves and hover need flight
@dataclass
class TypedGraph: # G = (V, E, X, S, R). S and R live in the functions below.
nodes: dict # id -> canonical or raw action label
edges: list # (src_id, dst_id)
xtype: str
params: dict = field(default_factory=dict) # id -> {"needs_grounding": bool}
mapped: set = field(default_factory=set) # ids whose label was rewritten through SYNONYMS
diag: list = field(default_factory=list) # construction diagnostics D
def _norm(label):
return str(label).strip().lower().replace(" ", "_")
def construct(candidate):
"""Algorithm 1. Schema consistency, edge recovery, canonical actions, diagnostics. Never authorizes anything."""
diag, nodes, params, mapped = [], {}, {}, set()
for n in candidate.get("nodes", []):
nid = n.get("id")
lab = n.get("command", n.get("action", n.get("name"))) # equivalent field names
if nid is None or lab is None:
diag.append("malformed_node"); continue
lab = _norm(lab)
if lab not in A_IDX and lab in SYNONYMS:
lab = SYNONYMS[lab]; mapped.add(nid) # safely synonym mappable
nodes[nid] = lab
params[nid] = {"needs_grounding": bool(n.get("needs_grounding", False))}
edges = []
for e in candidate.get("edges", []):
s, t = (e if isinstance(e, (list, tuple)) else
(e.get("from", e.get("source", e.get("src"))), e.get("to", e.get("target", e.get("dst")))))
if s in nodes and t in nodes:
edges.append((s, t)) # resolved dependency edges only
else:
diag.append("bad_edge")
return TypedGraph(nodes, edges, _norm(candidate.get("execution_type", "")), params, mapped, diag)
def generations(g):
"""Longest path layering. Returns (list of generations, acyclic flag). Start time is the generation index."""
indeg = {n: 0 for n in g.nodes}
for s, t in g.edges:
indeg[t] += 1
layer = {n: 0 for n in g.nodes}
ready = [n for n, d in indeg.items() if d == 0]
seen = 0
while ready:
n = ready.pop(); seen += 1
for s, t in g.edges:
if s == n:
layer[t] = max(layer[t], layer[n] + 1); indeg[t] -= 1
if indeg[t] == 0:
ready.append(t)
if seen != len(g.nodes):
return [], False
out = [[] for _ in range(max(layer.values()) + 1)] if layer else []
for n in sorted(g.nodes):
out[layer[n]].append(n)
return out, True
def infer_type(g):
gens, ok = generations(g)
if len(g.nodes) == 1 and not g.edges:
return "single"
if not g.edges:
return "parallel"
if ok and all(len(x) == 1 for x in gens):
return "sequential"
return "mixed"
def components(g):
parent = {n: n for n in g.nodes}
def find(x):
while parent[x] != x:
parent[x] = parent[parent[x]]; x = parent[x]
return x
for s, t in g.edges:
parent[find(s)] = find(t)
comps = {}
for n in g.nodes:
comps.setdefault(find(n), []).append(n)
return sorted(comps.values(), key=lambda c: min(c))
def violations(g, s0="grounded"):
"""The predicate set R. Returns a set of violation codes. Each predicate is a check on G and s0 (Eq. 6a)."""
v = set()
if g.diag:
v.add("bad_structure")
if any(a not in A_IDX for a in g.nodes.values()):
v.add("illegal_command") # legality against A
if g.mapped:
v.add("synonym_label") # legal only after the mapping
if any(p["needs_grounding"] for p in g.params.values()):
v.add("missing_grounding") # routed to clarify
gens, acyclic = generations(g)
if not acyclic:
v.add("cycle"); return v # nothing else is meaningful
if g.xtype not in EXEC_TYPES or g.xtype != infer_type(g):
v.add("incorrect_execution_type")
if g.xtype == "single" and (len(g.nodes) != 1 or g.edges):
v.add("single_sanity")
sig = {}
for n, a in g.nodes.items(): # structural twins are redundant
key = (a, frozenset(s for s, t in g.edges if t == n), frozenset(t for s, t in g.edges if s == n))
sig.setdefault(key, []).append(n)
if any(len(x) > 1 for x in sig.values()):
v.add("redundant_nodes")
if infer_type(g) in ("sequential", "mixed") and len(components(g)) > 1:
v.add("disconnected_graph")
state = s0
for gen in gens:
acts = [g.nodes[n] for n in gen]
if any(frozenset((a, b)) in CONFLICTS for a, b in itertools.combinations(acts, 2)):
v.add("conflicting_commands") # same generation conflict pair
if all(a in A_IDX for a in acts):
nxt = state
for a in acts:
if delta(state, a) is None:
v.add("state_violation")
nxt = delta(state, a) or nxt
state = nxt
# Eqs. 3 and 5 hold by construction of the layering: tau_e(pred) = layer + 1 <= tau_s(succ).
return v
HARD = {"cycle", "conflicting_commands", "state_violation", "illegal_command", "bad_structure", "single_sanity"}
REPAIRABLE = {"redundant_nodes", "disconnected_graph", "incorrect_execution_type", "synonym_label"}
def repair_once(g):
"""One pass of the bounded repair policy P_repair. Only predefined recoverable violations are touched."""
v = violations(g)
if "redundant_nodes" in v:
seen, drop = {}, set()
for n in sorted(g.nodes):
key = (g.nodes[n], frozenset(s for s, t in g.edges if t == n), frozenset(t for s, t in g.edges if s == n))
(drop.add(n) if key in seen else seen.setdefault(key, n))
for n in drop:
del g.nodes[n]; g.params.pop(n, None)
g.edges = [(s, t) for s, t in g.edges if s not in drop and t not in drop]
if "disconnected_graph" in v:
comps = components(g)
for c1, c2 in zip(comps, comps[1:]): # inferable missing edges
sinks = [n for n in c1 if not any(s == n for s, t in g.edges)]
srcs = [n for n in c2 if not any(t == n for s, t in g.edges)]
g.edges += [(a, b) for a in sinks for b in srcs]
if "synonym_label" in v:
g.mapped = set()
g.xtype = infer_type(g) # relabel the execution type from topology
return g
def shield(candidate, s0="grounded"):
"""Algorithm 2. Returns (decision, authorized graph or None, violation codes)."""
g = construct(candidate)
v = violations(g, s0)
if not v:
return "accept", g, v
if v & HARD:
return "reject", None, v
if v <= REPAIRABLE:
for _ in range(MAX_REPAIR_ITERS): # deterministic termination
g = repair_once(g)
if not violations(g, s0): # complete revalidation before dispatch
return "repair", g, v
return "reject", None, v
if "missing_grounding" in v and not (v - {"missing_grounding"}):
return "clarify", None, v
return "reject", None, v
def ged_normalized(ref, pred):
"""Eqs. 8a and 8b. Graphs are (labels, edge set over indices). Brute force is fine up to about 6 nodes.
DESIGN CHOICE: unit costs, node relabel 1, node insert or delete 1, edge insert or delete 1."""
(Vr, Er), (Vp, Ep) = ref, pred
m = max(len(Vr), len(Vp)); best = 10 ** 9
for perm in itertools.permutations(range(m)):
c = 0
for i in range(m):
j = perm[i]
if i < len(Vr) and j < len(Vp):
c += Vr[i] != Vp[j]
elif i < len(Vr) or j < len(Vp):
c += 1
c += len({(perm[a], perm[b]) for a, b in Er} ^ set(Ep))
best = min(best, c)
z = len(Vr) + len(Vp) + len(Er) + len(Ep)
return best / z if z else 0.0
# ----------------------------------------------------------------------------------------
# PART B. Synthetic commands and a stand in proposer (DESIGN CHOICE throughout)
# ----------------------------------------------------------------------------------------
VERBS = ["move", "go", "fly", "head", "slide", "drift"]
HELD = {"left": ["fly", "head"], "right": ["head", "drift"], "forward": ["drift", "slide"], "backward": ["slide", "fly"]}
# DESIGN CHOICE: every verb appears with every direction in training except the held out pairs, so the test
# asks for new verb and direction combinations, not new words. Without this the small model learns that
# "drift" means backward, which is a shortcut and not a fair test.
def _lateral(d, held):
return [f"{v} {d}" for v in VERBS if (v in HELD[d]) == held]
TRAIN_PH = {"takeoff": ["take off", "lift off"], "land": ["land", "touch down"],
"move_up": ["move up", "ascend", "rise", "go higher", "fly up"], "move_down": ["move down", "descend", "sink", "go lower", "fly down"],
"move_left": _lateral("left", False), "move_right": _lateral("right", False),
"move_forward": _lateral("forward", False), "move_backward": _lateral("backward", False),
"hover": ["hover", "hold position"], "stop": ["stop", "halt"]}
TEST_PH = {"takeoff": ["take off", "lift off"], "land": ["land", "touch down"], # held out recombinations
"move_up": ["go up", "rise higher", "move higher", "ascend"], "move_down": ["go down", "sink lower", "move lower", "descend"],
"move_left": _lateral("left", True), "move_right": _lateral("right", True),
"move_forward": _lateral("forward", True), "move_backward": _lateral("backward", True),
"hover": ["hover", "hold position"], "stop": ["stop", "halt"]}
SEQ_TRAIN, PAR_TRAIN = ["then", "and then", "after that"], ["and at the same time", "while", "and simultaneously"]
SEQ_TEST, PAR_TEST = SEQ_TRAIN, PAR_TRAIN # DESIGN CHOICE: connectors are shared. Held out connectors sank a model this small.
MAX_NODES = 4
def sample_plan(rng, valid):
"""Stages of actions. Nodes in one stage run in parallel, consecutive stages are fully connected."""
n_stages = rng.choice([1, 2, 2, 3]); stages, total = [], 0
for si in range(n_stages):
size = 1 if total >= 3 else rng.choice([1, 1, 2])
size = min(size, MAX_NODES - total)
if valid:
if si == 0:
acts = ["takeoff"]
elif si == n_stages - 1 and rng.random() < 0.3:
acts = ["land"]
else:
pool = MOVES + ["hover"]; acts = []
while len(acts) < size:
a = rng.choice(pool)
if a not in acts and all(frozenset((a, b)) not in CONFLICTS for b in acts):
acts.append(a)
else:
acts = [rng.choice(ACTIONS) for _ in range(size)]
stages.append(acts); total += len(acts)
if total >= MAX_NODES:
break
return stages
def plan_to_graph(stages):
labels, stage_of = [], []
for si, st in enumerate(stages):
for a in st:
labels.append(a); stage_of.append(si)
edges = [(i, k) for i in range(len(labels)) for k in range(len(labels)) if stage_of[k] == stage_of[i] + 1]
return labels, edges
def graph_to_candidate(labels, edges, xtype):
return {"execution_type": xtype, "nodes": [{"id": f"n{i}", "command": a} for i, a in enumerate(labels)],
"edges": [{"from": f"n{i}", "to": f"n{k}"} for i, k in edges]}
def true_type(labels, edges):
g = construct(graph_to_candidate(labels, edges, "single")); return infer_type(g)
def realize(stages, rng, ph, seq, par):
parts = []
for st in stages:
s = rng.choice(ph[st[0]])
for a in st[1:]:
s += " " + rng.choice(par) + " " + rng.choice(ph[a])
parts.append(s)
out = parts[0]
for p in parts[1:]:
out += " " + rng.choice(seq) + " " + p
return out
def make_dataset(n, seed, split, valid_frac=0.6):
rng = random.Random(seed)
ph, seq, par = (TRAIN_PH, SEQ_TRAIN, PAR_TRAIN) if split == "train" else (TEST_PH, SEQ_TEST, PAR_TEST)
rows = []
for _ in range(n):
stages = sample_plan(rng, rng.random() < valid_frac)
labels, edges = plan_to_graph(stages)
rows.append((realize(stages, rng, ph, seq, par), labels, edges, true_type(labels, edges)))
return rows
class Vocab:
def __init__(self, texts):
words = sorted({w for t in texts for w in t.split()})
self.idx = {w: i + 2 for i, w in enumerate(words)} # 0 pad, 1 unknown
def __len__(self):
return len(self.idx) + 2
def encode(self, t, L=40):
ids = [self.idx.get(w, 1) for w in t.split()][:L]
return ids + [0] * (L - len(ids))
PAD_ACT = len(ACTIONS) # slot has no action
def tensorize(rows, vocab):
tok = torch.tensor([vocab.encode(r[0]) for r in rows])
act = torch.full((len(rows), MAX_NODES), PAD_ACT, dtype=torch.long)
rel = torch.zeros(len(rows), MAX_NODES, MAX_NODES)
typ = torch.tensor([EXEC_TYPES.index(r[3]) for r in rows])
for b, (_, labels, edges, _) in enumerate(rows):
for i, a in enumerate(labels):
act[b, i] = A_IDX[a]
for i, k in edges:
rel[b, i, k] = 1
return tok, act, rel, typ
class Proposer(nn.Module):
"""DESIGN CHOICE stand in for the LLM. Transformer encoder, learned slot queries (DETR style),
an action head per slot, a pairwise precedence head, and an execution type head."""
def __init__(self, vocab, d=64, heads=4, layers=2):
super().__init__()
self.emb = nn.Embedding(vocab, d, padding_idx=0); self.pos = nn.Embedding(64, d)
self.enc = nn.TransformerEncoder(nn.TransformerEncoderLayer(d, heads, 4 * d, 0.1, batch_first=True, norm_first=True), layers)
self.slot_q = nn.Parameter(torch.randn(MAX_NODES, d) * 0.02)
self.cross = nn.MultiheadAttention(d, heads, batch_first=True)
self.act_head = nn.Linear(d, len(ACTIONS) + 1)
self.rel_head = nn.Sequential(nn.Linear(3 * d, d), nn.GELU(), nn.Linear(d, 1))
self.type_head = nn.Linear(d, len(EXEC_TYPES))
def forward(self, tok):
pad = tok == 0
h = self.enc(self.emb(tok) + self.pos(torch.arange(tok.size(1), device=tok.device)), src_key_padding_mask=pad)
q = self.slot_q.unsqueeze(0).expand(tok.size(0), -1, -1)
s, _ = self.cross(q, h, h, key_padding_mask=pad)
si, sk = s.unsqueeze(2).expand(-1, -1, MAX_NODES, -1), s.unsqueeze(1).expand(-1, MAX_NODES, -1, -1)
rel = self.rel_head(torch.cat([si, sk, si * sk], -1)).squeeze(-1)
pooled = (h * (~pad).unsqueeze(-1)).sum(1) / (~pad).sum(1, keepdim=True)
return self.act_head(s), rel, self.type_head(pooled)
def proposer_loss(out, act, rel, typ):
"""L = CE(actions) + BCE(edges on real upper triangle pairs) + CE(type). Equal weights are a DESIGN CHOICE."""
a_logit, r_logit, t_logit = out
real = (act != PAD_ACT)
pair = (real.unsqueeze(2) & real.unsqueeze(1)) & torch.triu(torch.ones(MAX_NODES, MAX_NODES, dtype=torch.bool), 1)
l_rel = F.binary_cross_entropy_with_logits(r_logit[pair], rel[pair]) if pair.any() else r_logit.sum() * 0
return F.cross_entropy(a_logit.reshape(-1, a_logit.size(-1)), act.reshape(-1)) + l_rel + F.cross_entropy(t_logit, typ)
def train(rows, vocab, epochs=20, bs=64, lr=3e-3, seed=0):
torch.manual_seed(seed)
model = Proposer(len(vocab)); opt = torch.optim.AdamW(model.parameters(), lr=lr, weight_decay=1e-2)
tok, act, rel, typ = tensorize(rows, vocab); n = len(rows); hist = []
sched = torch.optim.lr_scheduler.OneCycleLR(opt, max_lr=lr, total_steps=epochs * math.ceil(n / bs))
for ep in range(epochs):
model.train(); perm = torch.randperm(n); tot = 0.0
for i in range(0, n, bs):
j = perm[i:i + bs]
loss = proposer_loss(model(tok[j]), act[j], rel[j], typ[j])
opt.zero_grad(); loss.backward(); opt.step(); sched.step(); tot += loss.item() * len(j)
hist.append(tot / n)
return model, hist
@torch.no_grad()
def propose(model, vocab, texts):
"""Decode proposals into the JSON like candidate structure that Algorithm 1 consumes."""
model.eval(); tok = torch.tensor([vocab.encode(t) for t in texts])
a, r, t = model(tok); out = []
for b in range(len(texts)):
labels = [ACTIONS[i] for i in a[b].argmax(-1).tolist() if i != PAD_ACT]
k = len(labels)
edges = [(i, j) for i in range(k) for j in range(i + 1, k) if r[b, i, j] > 0]
out.append((labels, edges, EXEC_TYPES[t[b].argmax().item()]))
return out
def evaluate(model, vocab, rows):
props = propose(model, vocab, [r[0] for r in rows])
n = len(rows); exact = type_ok = acc_ok = wrong = wrong_passed = 0; ged = 0.0; dec = {"accept": 0, "repair": 0, "reject": 0, "clarify": 0}
oracle_agree = 0; t_shield = 0.0; wrong_valid = wrong_valid_passed = 0
for (txt, labels, edges, xt), (pl, pe, px) in zip(rows, props):
same = pl == labels and set(pe) == set(edges)
exact += same; type_ok += px == xt
ged += ged_normalized((labels, set(edges)), (pl, set(pe)))
t0 = time.perf_counter(); d, g, v = shield(graph_to_candidate(pl, pe, px)); t_shield += time.perf_counter() - t0
d_true, _, _ = shield(graph_to_candidate(labels, edges, xt))
dec[d] += 1; oracle_agree += (d == d_true)
if not same:
wrong += 1; wrong_passed += d in ("accept", "repair")
if d_true == "accept": # the user asked for something legal
wrong_valid += 1; wrong_valid_passed += d in ("accept", "repair")
return dict(n=n, exact=exact / n, type_acc=type_ok / n, mean_nged=ged / n, decisions=dec, oracle_agree=oracle_agree / n,
wrong=wrong, wrong_passed=wrong_passed, wrong_valid=wrong_valid, wrong_valid_passed=wrong_valid_passed, shield_us=1e6 * t_shield / n)
# ----------------------------------------------------------------------------------------
# Shield unit test that mirrors the structure of Table 21. Faults are built from the same taxonomy
# the rules encode, so a perfect score is expected. It checks the implementation, not coverage of reality.
# ----------------------------------------------------------------------------------------
def fault_cases(rng, per_class=60):
def seq(k):
acts = ["takeoff"] + [rng.choice(MOVES + ["hover"]) for _ in range(k - 1)]
return acts, [(i, i + 1) for i in range(k - 1)]
cases = {c: [] for c in ["cycle", "conflicting_commands", "redundant_nodes", "disconnected_graph",
"state_violation", "illegal_command", "incorrect_execution_type"]}
for _ in range(per_class):
k = rng.choice([3, 4]); a, e = seq(k)
while True: # flying start, no takeoff, no axis conflicts
d = [rng.choice(MOVES + ["hover"]) for _ in range(k)]
if all(frozenset(p) not in CONFLICTS for p in itertools.combinations(d, 2)):
break
cases["cycle"].append((graph_to_candidate(a, e + [(k - 1, 0)], "sequential"), "grounded", "reject"))
ax = rng.choice([("move_up", "move_down"), ("move_left", "move_right"), ("move_forward", "move_backward")])
cases["conflicting_commands"].append((graph_to_candidate(["takeoff", ax[0], ax[1]], [(0, 1), (0, 2)], "mixed"), "grounded", "reject"))
cases["redundant_nodes"].append((graph_to_candidate(["takeoff", "move_up", "move_up"], [(0, 1), (0, 2)], "mixed"), "grounded", "repair"))
cases["disconnected_graph"].append((graph_to_candidate(d, [x for x in e if x != (1, 2)], "sequential"), "flying", "repair"))
cases["state_violation"].append((graph_to_candidate(["takeoff", "land", rng.choice(MOVES)], [(0, 1), (1, 2)], "sequential"), "grounded", "reject"))
cand = graph_to_candidate(a, e, "sequential"); cand["nodes"][1]["command"] = rng.choice(["barrel_roll", "return_to_base"])
cases["illegal_command"].append((cand, "grounded", "reject"))
cases["incorrect_execution_type"].append((graph_to_candidate(a, e, "parallel"), "grounded", "repair"))
return cases
def unit_tests(seed=0, per_class=60):
rng = random.Random(seed); res = {}
for name, cs in fault_cases(rng, per_class).items():
det = sum(name in shield(c, s0)[2] for c, s0, _ in cs)
right = sum(shield(c, s0)[0] == want for c, s0, want in cs)
res[name] = (det / len(cs), right / len(cs))
valid = [plan_to_graph(sample_plan(rng, True)) for _ in range(1000)]
false_reject = sum(shield(graph_to_candidate(l, e, true_type(l, e)))[0] == "reject" for l, e in valid) / 1000
# Algorithm 1 schema recovery, clarify path, synonym repair
alias = {"execution_type": "sequential", "nodes": [{"id": "a", "action": "Take Off"}, {"id": "b", "name": "ascend"}],
"edges": [["a", "b"]]}
ground = {"execution_type": "single", "nodes": [{"id": "a", "command": "move_forward", "needs_grounding": True}], "edges": []}
extras = dict(alias=shield(alias)[0], clarify=shield(ground, "flying")[0])
return res, false_reject, extras
def smoke_test():
warnings.filterwarnings("ignore", message="enable_nested_tensor")
t0 = time.time(); torch.manual_seed(0)
# 1. unit tests of the symbolic shield
res, false_reject, extras = unit_tests()
for k, (det, right) in res.items():
print(f"shield {k:<26} detected {det:.2f} right decision {right:.2f}")
print(f"shield false rejection on 1000 valid graphs {false_reject:.3f} alias recovery {extras['alias']} grounding gap {extras['clarify']}")
assert all(d == 1.0 and r == 1.0 for d, r in res.values()) and false_reject == 0.0
assert extras == {"alias": "repair", "clarify": "clarify"}
assert abs(ged_normalized((["a", "b"], {(0, 1)}), (["a", "b"], {(0, 1)}))) < 1e-9
# 2. stand in proposer, trained on synthetic commands, tested on held out phrase recombinations
train_rows = make_dataset(4000, 1, "train"); test_rows = make_dataset(1500, 2, "test")
vocab = Vocab([r[0] for r in train_rows])
model, hist = train(train_rows, vocab)
print(f"proposer parameters {sum(p.numel() for p in model.parameters())/1e3:.0f}k loss {hist[0]:.3f} to {hist[-1]:.3f}")
assert hist[-1] < hist[0]
m = evaluate(model, vocab, test_rows)
print(f"proposer graph exact match {m['exact']:.3f} type accuracy {m['type_acc']:.3f} mean nGED {m['mean_nged']:.4f}")
print(f"pipeline shield decisions {m['decisions']} agrees with decision on the true graph {m['oracle_agree']:.3f}")
print(f"pipeline wrong proposals {m['wrong']} of {m['n']}, of which accepted or repaired {m['wrong_passed']}")
print(f"pipeline wrong proposals for legal requests {m['wrong_valid']}, of which the shield let through {m['wrong_valid_passed']} shield time {m['shield_us']:.0f} microseconds per graph")
# 3. the point of the exercise, a legal but wrong graph sails through
wrong_way = graph_to_candidate(["takeoff", "move_left"], [(0, 1)], "sequential") # user said move right
print("legal but wrong graph (asked for right, got left) ->", shield(wrong_way)[0])
assert shield(wrong_way)[0] == "accept"
print(f"smoke test passed in {time.time() - t0:.0f} s")
if __name__ == "__main__":
smoke_test()
The smoke test runs three things. It first puts the shield through a unit test that mirrors the structure of the paper’s validation table, with 60 faulty graphs for each of seven violation classes plus 1000 valid graphs. It then trains the proposer on 4000 synthetic commands and tests it on 1500 held out ones whose verb and direction pairings never appeared in training. Finally it passes the proposals through the shield and compares the verdicts with those on the true graphs. Here is the output from one run on a CPU.
Start with the shield rows. Every class is detected and every decision is the intended one, and none of the 1000 valid graphs is rejected. That result is expected, because we wrote the faults from the same list as the rules. It shows the code matches the specification and nothing more, which is also the honest reading of the paper’s 2100 graph result. Two extra checks cover paths the paper describes but does not tabulate. A graph with alternate field names and a synonym (“ascend”, “Take Off”) is normalized and repaired, and a node that needs grounding is routed to clarify and not authorized.
The proposer rows are a plumbing check. It has 137 thousand parameters and reaches 97.7 percent exact graph match on the held out pairings, with a mean normalized edit distance of 0.018. That says very little about language models. The grammar is tiny, the vocabulary is about forty words, and an earlier version with unseen connector phrases collapsed to 40 percent exact match. We kept the connectors shared for that reason, and we mention it because it is a small echo of the paper’s stress test, where paraphrases hurt far more than noise did.
The pipeline rows contain a result we did not plan. The shield rejected all 34 wrong proposals, including the nine that answered legal requests. In this toy the proposer’s mistakes were of the rule breaking kind, such as an extra takeoff or a duplicated node, and those are exactly what the rules catch. A failure of that kind is safe, and it is also a refusal of a good command. That outcome reflects the model’s error types and not a general property of shields. The last line makes the general point with a constructed case. A graph that takes off and moves left, when the speaker asked for right, is accepted, because every predicate holds. This is the blind spot the paper’s stress test hints at, where 55 of 114 graphs were risky and the rules fired 11 times.
On speed, the shield took about 40 microseconds per graph here, including construction, which is the same order of magnitude as the paper’s 0.0261 ms. The hardware, language and measurement method differ, so read that as agreement about the scale and not a match. The cost of deterministic checking on graphs this small is negligible next to a language model call.
What this adds up to
The core achievement is a working demonstration of a simple division of labor. A language model turns spoken sentences into typed execution graphs. Expert knowledge, which is ten actions, three conflict axes and a small flight state machine, decides what counts as legal, and a deterministic shield decides what reaches the drone. On a 1000 command set the best structuring model reached 89.6 percent, the validator flagged every one of 2100 faulty graphs, 12 pilot users completed 88.3 percent of their tasks, and 75 of 80 validated commands flew successfully indoors. Those are respectable numbers for a first system that is open about being task level and indoor.
The conceptual shift is about where trust lives. Much of the current excitement about language models in robotics treats them as planners that will get good enough to be trusted. This paper takes the opposite and more durable view. The model is allowed to be wrong, the question is what happens next, and the answer is a piece of code that a person can read end to end. Soundness with respect to written rules is a modest guarantee, but it is a guarantee, which is more than a model’s confidence score can offer. The architecture would still be sensible if the language model were replaced by a better one next year.
The pattern should travel well. Any setting where a model proposes actions for a machine has the same three needs, a vocabulary of allowed actions, a notion of legal order and a list of things that must not happen together. Warehouse robots, lab automation, agricultural machinery, vehicle cabins and factory cells all have them, and in each case the rule layer is cheap, inspectable and independent of the model. The idea of repair driven by diagnostics also carries over, since the paper’s numbers show that telling the model exactly which rule it broke helps, even if the margin over a blind retry is modest.
The limits are real, and the authors list most of them. The action set has ten entries. The test sets are small, the faults in the validation experiments were built from the rules being tested, and the network results come from a simulation. The stress test shows the shield passing many graphs that are legal and wrong, because legality is all it can judge. The pilot has no comparison condition, the indoor drone weighs under 55 grams, the delay from speech to action is about 2.5 seconds, and the stop command shares that delay. The baselines use different metrics from the language models, and two of the significance results depend on a five to one difference in sample size. None of this undermines the design, but it marks the distance between a convincing laboratory system and one that belongs near people.
The road ahead is sketched in the paper. The authors plan closed loop sensing so graphs can be revalidated and replanned during a mission, multimodal grounding for commands such as “the tree over there”, formal verification of graph constraints, expert reviewed and data assisted rule refinement, smaller task specific models that could run onboard, and hierarchical and multi vehicle missions. Our own wish list is short. A public code release would let others test the rules. A comparison against a joystick or a fixed phrase menu would show whether speech helps. A read back of the plan before flight would catch the legal but wrong case, and a local path for the stop command would make the most urgent word the fastest.
The paper ends where good safety engineering usually begins, with a clear statement of what each layer is responsible for. A language model can be wonderfully flexible about how people talk, as long as something plain and checkable stands between its words and the rotors.
Frequently asked questions
How does speech control of a drone with a large language model work?
The spoken command is transcribed to text, and a language model turns it into a directed graph in which each node is one of ten canonical drone actions and each edge says what must finish first. A rule based shield then checks the graph and returns accept, reject, repair or clarify. Only an accepted or repaired graph is sent to the drone controller.
Why not let the language model control the drone directly?
A fluent plan is not necessarily an executable one. A model can name an action the drone lacks, skip a precondition such as being airborne, or order steps wrongly, for example landing and then moving forward. The authors treat the model’s output as a proposal and keep execution authority in deterministic rules that a reviewer can read and test.
How accurate was the system in the paper?
Command comprehension reached up to 95.4 percent with GPT 4 Turbo on the simplest tier. Gemma2 9B structured 1000 commands into graphs with 89.6 percent overall accuracy, against 73.1 to 76.0 percent for a rule parser and two TF IDF classifiers. Mixed commands that combine ordering with concurrency were the hardest, at 72.1 percent for the best model.
What does the safety shield actually guarantee?
It guarantees soundness with respect to the encoded expert knowledge. An accepted graph satisfies the rules for command legality, acyclicity, flight state transitions, execution type and same generation conflicts. It does not guarantee that the graph matches what the speaker meant. A graph that swaps left for right is legal and passes every rule.
Is it fast enough for real time flight control?
No, and the paper does not claim it is. Mean end to end latency was 2480 ms with a 95th percentile of 4210 ms, mostly from language model inference. The rule checks take about 0.0261 ms per graph. The system is positioned as a task level command and mission planning layer, with stabilization and reactive control left to the flight controller.
Can this be used on a real drone today?
Not as published. The evidence comes from 12 pilot users, simulated network failures and 80 validated indoor commands on a 54.8 gram drone, using a vocabulary of ten actions. The paper lists no public code, and its data are available on request. Anyone flying a similar system outdoors must follow local aviation rules and keep a human able to take control.
Read the paper
The article is open access under a Creative Commons license in Engineering Applications of Artificial Intelligence. The paper states that its data are available on request and lists no public code repository, so the secondary button below leads to more analyses in the same category and not to a repository. Replace it with a repository link if one appears.
Choi, S. H., Kim, Z. C., and Buu, S. J. Speech guided unmanned aerial vehicle control system based on large language model structured execution graphs with expert rule constraints. Engineering Applications of Artificial Intelligence 184 (2026) 116349. DOI 10.1016/j.engappai.2026.116349. Open access under a CC BY NC 4.0 license.
This analysis is based on the published paper and an independent evaluation of its claims.
