k. Design Patterns Evaluation SFT1

Version 4.5 by Tjalling Haije on 2026/07/27 16:17

Approach

The goal for the human-machine teaming framework is to (among others) develop human-machine teaming patterns that can be used to use technology

Design patterns of interest: 

Approaches for evaluating DPs:

  1. Make new usecases based on how the FR actually did the TDP and IDP with the robots with claims and check those with the raw data. 
  2. Match the premade usecase UC04.3: Detailed indoor exploration with ANYMAL (USAR) with the incident that fits best, and evaluate the claims for the TDP and IDP of interest using results from that incident. 

Choosing approach 2 because it has been prepared better, and helps in improving the evaluation approach for SFT 2 by identifying what info we missed. 

Design Patterns Evaluation

Approach: 

  1. Figure out in which incidents TDP2 and IDP1 were used
  2. Check if the tech worked as required (or good enough to be of value) in those incidents
  3. Figure out which of those incidents correspond most with UC04.3: Detailed indoor exploration with ANYMAL (USAR)
  4. Try to fill in j. SFT1 evaluation templates from the raw results, which 
  5. Check for each of the claims the results for the connected measures > Evaluates UC
  6. Check for each of the pros/cons for the Design Pattern if they are validated > Evaluates Design Pattern
  7. Adjust the Design Patterns based on the results and/or their expected pros/cons.
  8. Update how the users should use the technology and/or the technology itself to achieve the objectives. 
  9. Note missing data that should be included in the evaluation next time.

1. Design Patterns per Incident

IncidentIncident DescriptionTech usedDesign Patterns
Incident 1

A collision caused by a major explosion has left casualties trapped in vehicles and created a hazardous chemical leak from a truck. The goal of the fire service is to rescue the casualties, identify and contain the chemical hazard, and safely manage the incident.

??
Incident 2

An explosion has caused the partial collapse of multiple industrial and office buildings, leaving an unknown number of people trapped or missing. The goal of the Urban Search and Rescue (USAR) teams and drone team is to assess the damage, locate and rescue victims, manage responder safety during a secondary collapse, and clear all affected buildings.

Incident 3 

A blast has left victims stranded on the upper floors of a high-rise residential building. The goal of the fire crew and drone team is to assess the building, locate and safely rescue the trapped occupants using ladder access, and clear the building.

?

?
Incident 4

A gas explosion and fire in a four-apartment residential building have left multiple occupants trapped and created a hazardous environment for firefighters conducting search, rescue, and firefighting operations. The goal of the fire crews is to extinguish the fire, rescue all occupants, manage firefighter safety during an emergency involving a downed rescuer, and clear the building.

?

?

Based on the above list the next steps zoom in on incident 2.

2. Incident 2 tech evaluation

TechHow was it usedIssuesDid it work sufficiently?
ASTRAL RTK Drone (outdoor)Baseline & Run 1-3: The drone team conducts an aerial assessment of the collapsed buildings, identifying structural risks, access points, and possible victim locations. 

Run 1 USAR 1: Drone had connection issues which delayed its deployment.

Warning

TODO: expand these issues with questionnaire answers

Yes
OWL Indoor Drone 

Run 2 USAR 2: used to explore building and search for objects and victims 

 

Run 2 USAR 1: OWL missed victim and entire floor. Object detection initially disabled. 

Warning

TODO: expand these issues with questionnaire answers

Unreliable
ANYmal Robot with Robot Arm (for ANYmal) 

Run 1 USAR 1: remove beam blocking door

Run 3 USAR 1: Used to close a valve

 

WarningTODO: expand these issues with questionnaire answers

 

Yes
ANYmal Robot with SNAKE Run 3 USAR 1: Used to search for victims in a narrow spaceWarningTODO: expand these issues with questionnaire answers?

3. Incident 2 way of working

Run & teamHow did the first responders use the techSimilar to what usecase
Run 1, USAR team 1

The outdoor drone provides the Team Leader with an initial assessment of the incident scene. The Team Leader uses this information to deploy USAR teams and robotic assets; ANYmal supports by investigating and clearing hazards, after which the USAR teams gain access, stabilize the structure, locate victims, and perform rescues while continuously reporting progress to the Team Leader.

UC04.3: Detailed indoor exploration with ANYMAL (USAR) 
Although the robot didn't explore inside.
Run 2, USAR team 2

The OWL indoor drone first explores the collapsed structure and streams its findings to the Team Leader. Based on this reconnaissance, the USAR team enters, searches for victims, responds to emergencies, and conducts rescues while the Team Leader maintains situational awareness through the digital support system.

UC04.3: Detailed indoor exploration with ANYMAL (USAR) 
Run 3, USAR team 1

The outdoor drone assesses the scene, SNAKE inspects the interior, and ANYmal handles hazardous intervention tasks. Using the information gathered by these robotic systems, the USAR team enters, locates victims, and performs the rescue while the Team Leader maintains overall situational awareness and coordination.

UC04.3: Detailed indoor exploration with ANYMAL (USAR) 
Although the robot didn't explore inside.

4. Raw results

Results for Incident 2, Run 3, USAR team 1 VS baseline USAR team 1

Analysis of questionnaire results for members with ID: UT13, UT11, UT12, UTL1

ScenarioRoleParticipants
BaselineUSAR Team Member3
BaselineUSAR Team Leader1
Run 3USAR Team Member3
Run 3USAR Team Leader1
Run 1USAR Team Member3
Run 1USAR Team Leader1

1785167968429-388.png

1785167981571-844.png

1785167994552-876.png

These results show that in run 1 workload, trust and situation awareness were significantly lower compared to the baseline, with decision making and safety confidence differing only slightly from the baseline. 
Run 3 showed an improvement in decision making (specified as right info was available and decisions were timely) compared to the baseline and run 1.

On average both runs had a lower workload (-25%), trust (-33%), situation awareness (-11%), slightly lower safety confidence (-4%) and improved decision making (+8%).

Note: Run 3 also functioned as a demonstration, which increased speed of decisions and the scenario as a whole. Also, participants performed incident 2 in 3 runs in the same teamcomposition and building layout, such that participants giving them an unfair advantage. These factors might have influenced the results.

5. Results connected to Design Pattern measures

Note: The measures are defined in UC04.3: Detailed indoor exploration with ANYMAL (USAR)  and x. Measures 

MeasureResultsConnected ClaimClaim proven?
Safe path chosenNo FR path data available.CL01 Improve safety for responders Unknown
Decision qualityCan we use decision making questions from questionnaire for this?  CL01 Improve safety for responders Unknown
Near-incident and incident reportsNo incidents in baseline or run. 
No near-incidents data available.
CL01 Improve safety for responders Unknown
Reported hazards, victims, layout correctness, based on robot intel before FR entryNo ground truth data available.CL02 Increased SAUnknown
How optimal was the FR pathNo FR path or ground truth data available.CL02 Increased SAUnknown
Reported hazards, victims, layout correctness, after FR entryNo FR path or ground truth data available.CL02 Increased SAUnknown
Workload: e.g. NASA-TLXQuestionnaire indicates increased workload of 25%. But not literature reviewed questionnaire and runs not clean. CL03 improved mission effectivenessIndication says strong no, but data unreliable.
Appropriate trust in technology: e.g. trust surveyCan we use trust questions from questionnaire for this? They did not target robots specifically.CL04 acceptable workloadUnknown
Decision confidence: e.g. Retrospective decision quality ratingCan we use decision making questions from questionnaire for this?  CL04 acceptable workloadUnknown
Trust in SAQuestionnaire indicates decreased trust of -33%. But not literature reviewed questionnaire and runs not cleanCL05 Trustworthy SAIndication says strong no, but data unreliable.
Health indicators such as HR, temperature, etc.No health data available.CL06 Improved FR healthUnknown
Health issues tackled in acceptable time?CL06 Improved FR healthUnknown
Total mission completion time?CL07 Degraded mission efficiencyUnknown
First responder time inside?CL07 Degraded mission efficiencyUnknown

(Reliable) data missing to properly evaluate usecase and as such team design pattern. 

5. Pros and cons for TDP

TDP2 — Robots First, Then Humans Protocol 

Pro / ConResultsPro / Con proven?