Test CFT4 and the coming IFT's
1. Introduction
This experiment validates multiple human-machine teaming technologies in Urban Search and Rescue (USAR) operations. Four operational modules simulate a full operational storyline across two days: wide area assessment, full-area reconnaissance with health monitoring, indoor drone-assisted search, and precision inspection in confined spaces.
The modules test system functions from five use cases, and aim to quantify effects on safety, situation awareness (SA), physical workload, mission effectiveness, and decision-making quality. Results will be compared against expected performance without these technologies, based on either baseline team data, observer input, or known solution benchmarks.
2. Method
2.1 Participants
Approximately 24–30 international first responders, organized in teams. Each team rotates across the four modules. Roles include responders, team leaders, drone/robot operators, analysts, medics, and safety officers.
2.2 Experimental Design
A within-subject design is used where all teams go through the four modules. Performance is compared across modules and against predefined baseline criteria. Observers collect data in real time; surveys and biometric data are used to validate subjective and objective measurements.
2.3 Tasks (Per Module)
Module 1 – Wide Area Assessment
Use Case: UC03.0
Scenario: Teams arrive at a simulated disaster zone. Structures are unstable. Drone support is requested for external mapping and hazard detection.
Tested Functions:
- Drone feed provides real-time visuals to field teams and command
- Zoom-ins allow inspection of rooftops and entry points
- Footage used to mark safe approach routes
Measured Claims:
- CL1: Improved external SA
- CL2: Safer movement planning
- CL3: Faster planning cycle
- CL4: Reduced mental workload for recon
- CL5: Improved coordination (shared SA)
Quantifiable Success Factors:
- ≥80% of hazards correctly marked on the map (based on preset dummy hazards)
- ≥90% agreement in SA between team and command (map match)
- Average planning time ≤ 10 minutes from drone launch
- NASA-TLX workload score ≤ 50 (moderate) for command roles
How to Measure:
- Observer logs & stopwatch for planning time
- Map test: Compare team-drawn vs. actual map (SAGAT-lite)
- Count number of correctly identified hazards from drone feed
- NASA-TLX filled by drone operator and team lead
- Post-module survey: "How useful was the drone in forming your plan?" (1–5)
Module 2 – Health Monitoring & Reconnaissance
Use Cases: UC01.1 (Fire) and UC01.2 (USAR)
Scenario: Team performs full-area recon. Wearables measure heart rate, hydration, and simulated gas exposure. Simulated fatigue and alerts escalate to medics or team leads.
Tested Functions:
- Alerts for fatigue/gas exposure
- Remote dashboard monitoring by safety officer
- Escalation protocols for health interventions
- Logging and after-action review
Measured Claims:
- CL1–CL2: Prevent overexertion and increase responder awareness
- CL3–CL4: Enable remote intervention and informed medical decision
- CL5: Enable better rotation/rest planning
- CL6: Debrief uses health logs
- CL7: Improve mission success
Quantifiable Success Factors:
- ≥90% of health alerts acknowledged within 1 minute
- ≥80% of interventions judged "timely" in AAR interviews
- ≥50% of teams adjust tactics or rest cycles based on health data
- ≥1 health-based lesson identified per team in debrief
- ≤2 simulated incidents due to unmanaged fatigue/gas exposure
How to Measure:
- Log alert timings vs. response time
- Observer notes + medic reports on intervention
- Exit survey: "Did alerts help prevent fatigue/injury?"
- Use of wearable dashboard during debrief (Yes/No)
- NASA-TLX for responders
Module 3 – Indoor Drone Search (Barracks)
Use Cases: UC02.1 and UC02.2
Scenario: Collapsed barracks building. Indoor drone used for autonomous scan. Analyst tags victims, hazards, and updates C3I map. Drone does close inspection on request.
Tested Functions:
- Pre-entry thermal scan
- Hazard/victim detection
- Analyst-supported interpretation and tagging
- Entry planning based on drone data
Measured Claims:
- CL1: Heightened SA before entry
- CL2: Increased safety (less exposure)
- CL3–CL5: Faster, more accurate victim detection
- CL6: Trust in drone data
- CL7: Increased mission efficiency
Quantifiable Success Factors:
- ≥90% of dummy victims detected by drone+analyst
- ≥2 new hazards marked per team from drone feed
- Average time-to-first victim ≤ 3 minutes
- ≥80% of responders rate drone info as “trustworthy” (score ≥4/5)
- ≤1 injury due to unknown hazard in follow-up entry
How to Measure:
- Victim tags placed in known positions for ground truth
- Observer logs: detection times and analyst confirmations
- Team SA quiz: "How many victims? Where were they located?"
- Trust survey: “I would act on this drone data” (1–5)
- Entry path compared to drone hazard map
Module 4 – Precision Inspection with ANYMAL/SNAKE
Use Case: UC04.0
Scenario: Teams reach unstable voids. Robots are deployed to inspect inaccessible areas. SNAKE arm is used to look into cracks. Results update team maps and entry plans.
Tested Functions:
- Autonomous or manual ANYMAL movement
- Void inspection using flexible arm
- Victim/hazard confirmation
- Decision-making based on robot visuals
Measured Claims:
- CL1: Access without risk
- CL2: Detection in confined space
- CL3: Safer routing
- CL4: Trust in robot-assessed visuals
- CL5: Faster room clearing
Quantifiable Success Factors:
- ≥2 hazards or victims confirmed via SNAKE per team
- ≥80% of voids scanned without human entry
- ≥70% of teams adjust route based on robot findings
- ≥80% of participants rate robot visuals as “clear and usable”
- Average inspection time ≤ 8 minutes per room
How to Measure:
- Observer log: robot path vs. human path
- Detection log compared to known hidden items
- Survey: “Did robot findings improve your plan?” (Yes/No)
- Video review of time-per-room
- Trust in visuals scale (1–5)
2.4 Measures
This section describes how each claim will be measured during each module, using a combination of objective logging, observer annotations, post-task surveys, and scenario-based evaluation.
Module 1 – Wide Area Assessment (UC03.0)
| Claim | Metric | Method/Tool | Success Threshold |
|---|---|---|---|
| CL1 – Improved SA | Number of hazards correctly identified on team maps | SAGAT-lite: Pre/post map-drawing task + verbal hazard recall | ≥80% match with ground-truth hazard list |
| CL2 – Safer planning | Number of hazard zones avoided during later entry | Observer logs cross-referenced with hazard map | 100% of marked hazards avoided |
| CL3 – Faster planning | Time from drone launch to team briefing | Stopwatch & observer notes | ≤10 minutes total |
| CL4 – Reduced workload | Mental workload score of command & drone operator | NASA-TLX (short form) | ≤50 average score |
| CL5 – Shared SA | Consistency between team and command in map data | Comparison of annotations across roles | ≥90% agreement on key features |
Module 2 – Health Monitoring & Reconnaissance (UC01.1 / UC01.2)
| Claim | Metric | Method/Tool | Success Threshold |
|---|---|---|---|
| CL1 – Prevent overload | HR trend + alert timing vs. pause/extraction | Wearable logs + observer notes | ≥90% alerts followed by correct action within 1 minute |
| CL2 – Responder awareness | Survey response on self-adjustment | Post-task Likert: “The alert helped me act” | ≥80% rate 4 or 5 |
| CL3 – Remote escalation | Alert-to-medic contact time | System log + stopwatch | ≤1 minute average |
| CL4 – Medical support | Alignment of alerts with medical assessment | Medic forms + sensor log correlation | ≥80% concordance |
| CL5 – Operational planning | Number of rest/rotation decisions based on dashboard | Observer logs + team lead AAR | ≥50% of teams adapt plan |
| CL6 – AAR use of health data | Was biometric data used during debrief? | Debrief analysis | Yes, per team |
| CL7 – Mission effectiveness | Task time + incidents avoided | Stopwatch + incident log | Task time not slower than baseline; 0 uncontrolled fatigue/gas incidents |
Module 3 – Indoor Drone Search (UC02.1 / UC02.2)
| Claim | Metric | Method/Tool | Success Threshold |
|---|---|---|---|
| CL1 – Heightened SA | SA questionnaire + map task | Pre/post: victims, layout, hazard count | ≥80% correct recall post-drone |
| CL2 – Increased safety | Hazard zone avoidance rate | Observer vs. ground truth map | ≥90% of flagged areas avoided |
| CL3 – Faster victim detection | Time to first detection | Stopwatch from drone entry | ≤3 minutes |
| CL4 – Accuracy of detection | Victim detection rate | Drone log vs. planted victims | ≥90% detected |
| CL5 – Trust in results | Survey: “I trust the drone data for decision-making” | 1–5 Likert scale | ≥80% rate 4 or 5 |
| CL6 – Efficiency | Entry time after drone plan vs. without drone | Stopwatch; compare with baseline data | 10–20% faster planning phase |
Module 4 – Robot-Based Precision Inspection (UC04.0)
| Claim | Metric | Method/Tool | Success Threshold |
|---|---|---|---|
| CL1 – Extended reach | Percentage of voids explored by robot not human | Observer log + inspection plan | ≥80% of voids scanned by robot |
| CL2 – Detection in small spaces | Victim/hazard detection in hidden locations | Camera log vs. planted markers | ≥2 findings per team |
| CL3 – Safer routing | Route changes based on robot input | Pre/post plan comparison + observer notes | ≥70% of teams adapt plan |
| CL4 – Trust in visuals | Survey on clarity and trust in robot data | Likert: “The robot data was sufficient for decisions” | ≥80% rate 4 or 5 |
| CL5 – Room clearing speed | Time per room before vs. after robot scout | Stopwatch log | ≤8 minutes per room avg. |
2.5 Procedure
All modules follow a similar four-part procedure, tailored per use case.
General Daily Timeline
- 08:30 – 09:00: Morning briefing, safety, tech setup
- 09:00 – 12:00: First module rotation (two parallel teams)
- 13:00 – 16:00: Second module rotation (two parallel teams)
- 16:00 – 17:00: Shared after-action review
Each module runs with the following structure:
Per Module Procedure
Briefing (10–15 min)
- Explain objectives, scenario, roles, safety, success factors
- Introduce technology and expectations
Execution Phase (45–60 min)
- Scenario runs in real time
- Observer logs events, actions, communications
- System logs recorded (drone, robot, wearables)
Measurement Phase (15–20 min)
- Paper or tablet surveys: SA, trust, NASA-TLX
- Sensor data downloaded to central system
- Short interview or checklist with operator and team lead
Debrief (15–20 min)
- Team reflects on use of technology, decision-making
- Facilitator prompts discussion of claims (trust, effectiveness, awareness)
- Recorded notes for final reporting
For cross-checking performance without the tech, one team per module may be assigned a simplified "control" version of the scenario, using conventional tools only (where feasible).
2.6 Material
Each module requires scenario-specific equipment, environmental props, logging tools, and survey forms:
Common Materials (all modules)
- Observer logbooks (standardized per module)
- Stopwatch or time-tracking app
- Participant role badges and checklists
- Data collection station with tablets/laptops
- Printed Likert-scale surveys (SA, trust, workload)
- SAGAT-lite map templates
Module-Specific Materials
Module 1 – Wide Area Assessment
- Outdoor drones with RTK GPS and live zoom cameras
- Command screen with drone feed
- Large printed site maps with hazard zones (for scoring)
- Structural hazard props (collapsed façades, signs)
Module 2 – Health Monitoring
- Wearable sensors (HR, hydration, gas; real or simulated)
- Dashboard software for live feed + logging
- Incident trigger devices (e.g., CO2 canisters, alarms)
- Medic checklist sheets
- Alert simulation software (optional)
Module 3 – Indoor Drone Search
- Thermal indoor drone with autonomous mode
- C3I-compatible map annotation system
- Dummy victims with heat packs or QR markers
- Printed room layouts for SA testing
- Indoor hazard props (rubble, fake smoke, blocked doors)
Module 4 – ANYMAL and SNAKE
- ANYMAL robot (legged) and SNAKE articulated arm
- Confined space mockups (voids, crawlspaces, stairs)
- Hidden hazard/victim tags inside small cavities
- Robot operator station + external monitor
- Scenario map with route overlays