We leverage LLMs to augment LTL specifications in RQ2 and RQ3. To systematically assess the quality of these LLM-generated LTLs, we conducted a structured evaluation involving six domain experts with professional experience in both TAP-enabled smart home systems and temporal logic specifications. The questions of the survey include:
(1) Expertise validation (Q1-Q2): Assessing participants' technical proficiency and practical experience
(2) Specification quality assessment (Q3-Q5): Evaluating the practical applicability, understandability, and potential redundancy of the generated specifications
Survey Demo
Expertise Validation Section
Q1:
Describe implementation strategies for achieving multi-protocol interoperability (Zigbee/Wi-Fi/Bluetooth Mesh) in TAP-enabled smart home systems.
(Reference answer: Unified semantic abstraction layer, priority-based event scheduling, and edge-computing security framework. )
Q2:
Apply negation normal form conversion to the LTL formula ¬G(Fϕ∧(ψ U θ)).
(Reference answer: ¬G(Fϕ∧(ψUθ))≡ F(G¬ϕ∨(¬ψ R ¬θ))
Mutated Specification Quality Assessment Section
Q3: Practical Relevance Evaluation
How would you assess the practical relevance of mutated specifications in real-world smart home applications?
Please rate on a scale of 1-5, where:
1 = The specification offers negligible improvement potential for TAP smart home systems
5 = The specification effectively identifies vulnerabilities in TAP systems and substantially enhances application security
Q4: Understandability and Readability Assessment
How would you rate the understandability and readability of mutated specifications?
Please score based on comprehension time (1-5 scale):
1 = Requires >60 seconds to understand
5 = Comprehensible in <15 seconds
Consider logical structure, terminology clarity, and syntactic complexity.
Q5: Potential Redundancy Analysis
How would you evaluate the conciseness of mutated specifications?
Rate using this criteria (1-5 scale):
1 = Highly redundant (e.g., nested repetitions like G(G(...)), duplicate atomic propositions)
5 = Optimally concise (minimal/no redundancy, efficient logical structure)
Example: "G(a ∧ G a)" could be simplified to "G a"
Our study involves a total of 200 specifications across Q3-Q5. We randomly selected 100 LLM-generated specifications (with their original specs) for systematic evaluation. To enable comprehensive analysis, we developed AutoLTL - a lightweight automated mutation tool that implements randomized mutation operators (atomic proposition insertion/replacement and temporal operator insertion/replacement) to generate LTLs. The mutation logic in AutoLTL strictly adheres to the mutation definitions provided in the LLM prompts. We used it to generate an additional 100 mutated specifications for comparative assessment. A representative comparative example is shown below:
Illustrative example of Specification Quality Assessment
Context: Evaluation based on the following smart home ruleset and LTL specifications
IFTTT Rules:
IF Smoke Sensor.smoke = detected & Motion Detector.motion = inactive THEN Alarm.both & Mobile Phone.take photo & Sina Weibo.post
IF Washer Machine.WasherMode = heavy & Weather.weather = sunny THEN Sina Weibo.post
IF Clock.time = 20 & Motion Detector.motion = active & TV.MachineState = off THEN Sprinkler Controller.open & Sina Weibo.post
IF Temperature Sensor.temperature >= 10 & Temperature Sensor.temperature <= 15 & Sina Weibo.State = posting & Air Conditioner.HvacMode = off THEN Online Bank.transfer money
IF Light.SwitchState = off & Clock.time = 23 THEN Sina Weibo.post
Original Specification:
G ( Washer Machine.WasherMode != heavy U Clock.time >= 9 )
Here is the corresponding answer from a PhD candidate to LLM-generated specifications and AutoLTL-generated specifications, respectively:
LLM-generated specifications
AutoLTL-generated specifications
Table 5 in the paper presents aggregated comparison results derived from this evaluation process.