Articles Open Access CC BY 4.0

LLM-Agent-Style Automated Usability Testing on MiniWoB++: A Reproducible Chunked Full-Run with ReAct, Plan-Execute, and Self-Reflect Policies

Authors

Article Integrity Crossmark: Check for updates
Article Metrics
171 Article Views 132 PDF Downloads 0 Crossref Citations

Abstract

Automated usability testing can reduce the cost of repeatedly checking whether a web interface supports reliable and efficient task completion, but existing scripted tests are brittle and many agent evaluations report benchmark scores without translating failures into usability diagnostics. This study asks how three LLM-agent-style strategies—ReAct, Plan-Execute, and Self-Reflect—differ in effectiveness, efficiency, and failure modes when applied to MiniWoB++ tasks, and whether their logged traces can support actionable UI analysis. We conducted a controlled experimental benchmark on 130 MiniWoB++ web tasks, running each strategy once under a fixed seed with the same deterministic DOM-grounded controller, headless Chromium harness, 10-step limit, and 2.0 s episode budget, producing 390 episodes. We analyzed task success, steps, wall-clock time, interaction category, difficulty bins, and failure categories using paired per-task comparisons and descriptive aggregation. Plan-Execute achieved the highest success rate (14.6%, 19/130), compared with 10.8% (14/130) for both ReAct and Self-Reflect; its advantage was most evident in form/transaction and selection tasks, while all strategies performed similarly on simple click/button tasks and failed on drag/scroll tasks. Failure analysis showed that wrong outcomes and element-grounding errors were the dominant bottlenecks, indicating that explicit planning improves coverage only when target elements can be reliably grounded. The findings contribute a reproducible baseline and a usability-oriented failure taxonomy for automated web-agent testing, suggesting that future frameworks should prioritize semantic grounding, plan validation, richer action primitives, and designer-facing diagnostics.

References

1
[1] International Organization for Standardization, “ISO 9241-11:2018 Ergonomics of human-system interaction—Part 11: Usability: Definitions and concepts,” Geneva, Switzerland: International Organization for Standardization, 2018.
2
[2] J. Nielsen, Usability Engineering. San Diego, CA, USA: Academic Press, 1993.
3
[3] J. Brooke, “SUS: A ‘quick and dirty’ usability scale,” in Usability Evaluation in Industry, P. W. Jordan, B. Thomas, B. A. Weerdmeester, and I. L. McClelland, Eds. London, U.K.: Taylor & Francis, 1996, pp. 189–194.
4
[4] E. Z. Liu, K. Guu, P. Pasupat, T. Shi, and P. Liang, “Reinforcement learning on web interfaces using workflow-guided exploration,” in Proc. International Conference on Learning Representations (ICLR), 2018. [Online]. Available: https://arxiv.org/abs/1802.08802
5
[5] T. Shi, A. Karpathy, L. Fan, J. Hernandez, and P. Liang, “World of Bits: An open-domain platform for web-based agents,” in Proc. 34th International Conference on Machine Learning (ICML), Proceedings of Machine Learning Research, vol. 70, 2017, pp. 3135–3144. [Online]. Available: https://proceedings.mlr.press/v70/shi17a.html
6
[6] Farama Foundation, “miniwob-plusplus: MiniWoB++ environment for Gymnasium,” GitHub repository, 2023. [Online]. Available: https://github.com/Farama-Foundation/miniwob-plusplus
7
[7] S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao, “ReAct: Synergizing reasoning and acting in language models,” in Proc. International Conference on Learning Representations (ICLR), 2023. [Online]. Available: https://arxiv.org/abs/2210.03629
8
[8] N. Shinn, F. Cassano, A. Gopinath, K. R. Narasimhan, and S. Yao, “Reflexion: Language agents with verbal reinforcement learning,” in Advances in Neural Information Processing Systems (NeurIPS), 2023. [Online]. Available: https://arxiv.org/abs/2303.11366
9
[9] L. Wang, W. Xu, Y. Lan, Z. Hu, Y. Lan, R. K.-W. Lee, and E.-P. Lim, “Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models,” in Proc. 61st Annual Meeting of the Association for Computational Linguistics (ACL), 2023, pp. 2609–2634. https://doi.org/10.18653/v1/2023.acl-long.147
10
[10] B. Xu, Z. Peng, B. Lei, S. Mukherjee, Y. Liu, and D. Xu, “ReWOO: Decoupling reasoning from observations for efficient augmented language models,” arXiv:2305.18323, 2023. [Online]. Available: https://arxiv.org/abs/2305.18323
11
[11] S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. Narasimhan, “Tree of Thoughts: Deliberate problem solving with large language models,” in Advances in Neural Information Processing Systems (NeurIPS), 2023. [Online]. Available: https://arxiv.org/abs/2305.10601
12
[12] J. Wei et al., “Chain-of-thought prompting elicits reasoning in large language models,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 35, 2022, pp. 24824–24837.
13
[13] T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom, “Toolformer: Language models can teach themselves to use tools,” in Advances in Neural Information Processing Systems (NeurIPS), 2023. [Online]. Available: https://arxiv.org/abs/2302.04761
14
[14] S. Yao et al., “WebShop: Towards scalable real-world web interaction with grounded language agents,” in Advances in Neural Information Processing Systems (NeurIPS), 2022. [Online]. Available: https://arxiv.org/abs/2207.01206
15
[15] S. Zhou et al., “WebArena: A realistic web environment for building autonomous agents,” in Proc. International Conference on Learning Representations (ICLR), 2024. [Online]. Available: https://arxiv.org/abs/2307.13854
16
[16] T. Le Sellier De Chezelles et al., “The BrowserGym ecosystem for web agent research,” arXiv:2412.05467, 2024. [Online]. Available: https://arxiv.org/abs/2412.05467
17
[17] Q. Wu et al., “AutoGen: Enabling next-gen LLM applications via multi-agent conversation,” arXiv:2308.08155, 2023. [Online]. Available: https://arxiv.org/abs/2308.08155
18
[18] J. Yang, C. E. Jimenez, A. Wettig, K. Lieret, S. Yao, K. Narasimhan, and O. Press, “SWE-agent: Agent-computer interfaces enable automated software engineering,” in Advances in Neural Information Processing Systems (NeurIPS), 2024. [Online]. Available: https://arxiv.org/abs/2405.15793
19
[19] L. Ouyang et al., “Training language models to follow instructions with human feedback,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 35, 2022, pp. 27730–27744. [Online]. Available: https://arxiv.org/abs/2203.02155
20
[20] P. F. Christiano et al., “Deep reinforcement learning from human preferences,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 30, 2017. [Online]. Available: https://papers.nips.cc/paper/7017-deep-reinforcement-learning-from-human-preferences
21
[21] R. Nakano et al., “WebGPT: Browser-assisted question-answering with human feedback,” arXiv:2112.09332, 2021. [Online]. Available: https://arxiv.org/abs/2112.09332
22
[22] J. Y. Koh et al., “VisualWebArena: Evaluating multimodal agents on realistic visual web tasks,” in Proc. 62nd Annual Meeting of the Association for Computational Linguistics (ACL), 2024. https://doi.org/10.18653/v1/2024.acl-long.50

Author Biography

Qi Xin

University of Pittsburgh United States

Management Information Systems

License

Downloads

Download data is not yet available.

Download Article

Downloads

Article Tools

View Issue

Citation Impact

Crossref Cited-by

0
0 citations

Counted from citation links between Crossref-registered works.

Last checked 24 Aug 2026.

How to Cite

Xin, Q. (2026). LLM-Agent-Style Automated Usability Testing on MiniWoB++: A Reproducible Chunked Full-Run with ReAct, Plan-Execute, and Self-Reflect Policies. Blockchain, Artificial Intelligence, and Future Research, 2(1), 56–82. https://doi.org/10.70211/bafr.v2i1.384