Journal of Advances in Shell Programming Review Article

Shell-Based Agents as Execution Systems: A Survey of Command Generation, Verification, Security, State, and Recovery

  1. Jyoti Dabass Golden Gate University
  2. Bhupender Singh Dabass District Bar Association

Abstract

Shell environments are increasingly becoming interfaces through which AI agents interact with software and operating systems. Unlike conventional command generation systems, shell-based agents can interpret task objectives, plan multi-step actions, invoke terminal tools, observe system state, and adapt subsequent actions. These capabilities support software development, system administration, experimentation, and automated operations, but also introduce reliability and security challenges that cannot be assessed through command generation alone. This survey reviews 57 studies published or publicly available between 2024 and 2026 on shell scripting, terminal-based agents, agentic workflows, execution environments, security, developer tools, and related evaluation methods. The reviewed studies were analyzed using structured evidence records covering research objectives, technical approaches, execution environments, evaluation strategies, findings, and limitations. The review examines shell-based agents as end-to-end execution systems, considering task interpretation, planning, command construction, execution, observation, state management, verification, and recovery. Findings indicate that syntax-aware generation, structured task representations, additional training signals, and multi-agent approaches can improve specific aspects of shell automation. However, reliable and secure execution also depends on tool interfaces, permission models, isolation mechanisms, state representations, approval processes, and recovery strategies. Security studies further indicate that unsafe behavior can originate from contextual sources such as logs, repository files, tool descriptions, external instructions, and agent skills. The review identifies a lack of standardized evaluation across functional correctness, security, efficiency, state consistency, and recovery. Future research should therefore emphasize trajectory-level evaluation, session-wide security analysis, fine-grained permissions, state-aware recovery, broader operating system coverage, and reproducible execution environments to enable reliable and controllable shell-based agents.

Keywords

References (57)

  1. Bin Y, Yuan X, Zeng H, Ye W, Shao W, Qian C, et al. Terminal agents: a survey of AI agents in command-line environments [Preprint]. 2026. arXiv:2608.20485.
  2. Jiang L, Huang S, Wu X, Li Y, Zhang D, Wei F. VisCodex: unified multimodal code generation via merging vision and coding models. International Conference on Learning Representations (ICLR). 2026.
  3. Liu Y, Zhang W, Yang Z, Zhang Z, Feng H, Wang X, et al. CARE: pre-execution command verification for shell-executing LLM agents [Preprint]. 2026. arXiv:2607.21642.
  4. Huang Z, Dundar R, Xie Y, Kallas K, Vasilakis N. Fractal: fault-tolerant shell-script distribution. 23rd USENIX Symposium on Networked Systems Design and Implementation (NSDI 26), Renton, WA, USA. 2026. p.2339-54.
  5. Wang S, Hou X, Liu Z, Zhao Y, Cheng X, Zou Q, et al. Demystifying and detecting agentic workflow injection vulnerabilities in GitHub Actions [Preprint]. 2026. arXiv:2605.07135.
  6. Yu L, Zhang J, Wang X, Yang L, Zhang F, Wang P, et al. Bash-commenter: leveraging syntax-aware preference optimization to reinforce large language model for bash code comment generation. Proc ACM Softw Eng. 2026;3:2117-40. doi:10.1145/3808102.
  7. Jiang W, Zhang X, Zhai J, Ma S, Shen C, Liu Y. False friends in the Shell: unveiling the emoticon semantic confusion in large language models. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), San Diego, CA, USA. 2026. p.26783-803. doi:10.18653/v1/2026.acl-long.1233.
  8. Pan S, Sun X, Zhang T, Liao D, Yang K, Xing Z. SkillGuard: a permission-centric framework for agent skill security [Preprint]. 2026. arXiv:2606.03024.
  9. Xu X, Saghir H, Wu Q, Côté MA, Wang T, Lakkaraju K, et al. The devil is in the interface: evaluating how tool architecture shapes coding agent behavior [Preprint]. 2026. arXiv:2608.11386.
  10. Mak H, Suresh S, Bhatnagar S, Wang B, Methani C, Munoz AG. Is Bash all you need? An empirical study of tool interfaces for enterprise digital worker agents [Preprint]. 2026. arXiv:2609.11999.
  11. Wang F, Shen G, Fu Y, Guo C, Cui Y, Chen Y. MABash: Multi-agent collaborative parsing for Bash command comment generation [Preprint]. 2026 Jun 1. doi:10.21203/rs.3.rs-9709428/v1.
  12. Zhang K, Chen Z, Feng X, Fang D, Zheng Y, Li Z, et al. Lightweight yet secure: secure scripting language generation via lightweight LLMs [Preprint]. 2026. arXiv:2601.06419.
  13. Du M, Luo Y, Banerjee S, Wojcik M, Popovic J, Cherukara MJ. Experiment automation agents: automating materials characterization with vision language model agents. NPJ Comput Mater. 2026. doi:10.1038/s41524-026-02213-8.
  14. Liu S, Tang X, Yang X, Lin L, Zhou B, Xiao W, et al. When the manual lies: a realistic benchmark to evaluate MCP poisoning attacks for LLM agents. 2026 29th International Conference on Computer Supported Cooperative Work in Design (CSCWD). 2026. p.2090-5.
  15. Yu L, Wang P, Xu J, Zhang J, Wang X, Ma J, et al. Towards robust and explainable Bash code generation with robustness-aware group relative policy optimization [Preprint]. 2026. arXiv:2606.27733.
  16. Rashidi M. The Balkanization of execution-security research for AI coding agents: isolation, access control, and time-of-check-to-time-of-use vulnerabilities [Preprint]. 2026. arXiv:2607.05743.
  17. Wu X, Zhao Z, Li Q, Li X, Shi Y, Adams B, et al. SkillShield: prompt-space security skills for LLM coding agents [Preprint]. 2026. arXiv:2608.25817.
  18. Alonso-Carracedo M, Fernandez-Boullon R, Celard P, Rodriguez-Martinez FJ, Otero-Cerdeira L. Automated grading of Linux/Bash examinations using large language models: a four-level cognitive taxonomy approach [Preprint]. 2026. arXiv:2607.02432.
  19. Li Z. Engineering a regular-type system for Unix and Linux shell [thesis]. Providence (RI): Brown University. 2026. Available from: https://atlas.cs.brown.edu/pdf/rt:brown:2026.pdf
  20. Hong J, Ascoli BG, Choi JD. ReCUBE: evaluating repository-level context utilization in code generation [Preprint]. 2026. arXiv:2603.25770.
  21. Xie Y, Lamprou E, Xia J, Vasilakis N. Incr: faster re-execution via bolt-on incrementalization. 20th USENIX Symposium on Operating Systems Design and Implementation (OSDI 26), Seattle, WA, USA. 2026. p.683-99.
  22. Liargkovas G, Jin D, Zhu TE, Liu D, Thompson AB, Narsipur A, et al. hS: speculative script reordering at subprocess granularity. 20th USENIX Symposium on Operating Systems Design and Implementation (OSDI 26), Seattle, WA, USA. 2026. p.665-81.
  23. Song L, Dai Y, Prabhu V, Zhang J, Shi T, Li L, et al. Coact-1: computer-using multi-agent system with coding actions. International Conference on Learning Representations (ICLR). 2026.
  24. Chen M, Zhang L, Feng Y, Wang X, Zhao W, Cao R, et al. SWE-Universe: scale real-world verifiable environments to millions [Preprint]. 2026. arXiv:2602.02361.
  25. Sladić M, Alibalić E, Valeros V, Catania C, Garcia S. AdvancedShelLM: a stateful multi-agent LLM honeypot for SSH deception [Preprint]. 2026. arXiv:2606.27990.
  26. Yeh CJ, Chang CH. PagePilot for web automation based on a multi-agent architecture. Proc Int AAAI Conf Web Soc Media. 2026;20(1):2623-36. doi:10.1609/icwsm.v20i1.42771.
  27. Khan HA, Khan I, Usmani KA, Ali M. Architecting next-generation shells: AI-driven command assistance and proactive safety mechanisms in the command-line. Int J Comput Appl. 2026;187:10-16. doi:10.5120/ijca4f98e99c1f94.
  28. Sladić M, Valeros V, Alibalić E, Garcia S. Improving LLM-based SSH honeypots through prompting and fine-tuning [Preprint]. 2026. arXiv:2608.18686.
  29. Jacobs J, Lapon J, Naessens V. Local LLMs for NL2Bash: a large-scale open-source model evaluation for Bash command generation. Workshop on LLM Assisted Security and Trust Exploration (LAST-X). 2026 Jan 20.
  30. Salviati U, De Gaspari F, Conti M, Mancini LV. ShellGames: speculative LLM-driven SSH deception [Preprint]. 2026. arXiv:2606.17986.
  31. Apostu AM, Preda A, Damir AD, Bolocan D, Ionescu RT, Croitoru I, et al. AutoMalDesc: large-scale script analysis for cyber threat research. Proc AAAI Conf Artif Intell. 2026;40(1):12-20. doi:10.1609/aaai.v40i1.36959.
  32. Kampourakis V, Chatzoglou E, Gkioulos V, Katsikas S. From conceptual scaffold to prototype: a standardized zonal architecture for Wi-Fi security training [Preprint]. 2026. arXiv:2605.07400.
  33. Curi de Miranda H, Lopes Roth Ferraz Á, Comin Sonaglio W, Alves Pereira Júnior L. A systematic security testing approach for InterUSS-based environments [Preprint]. 2026. arXiv:2605.
  34. Weng X. What you approve is what executes: consent integrity for black-box LLM agents [Preprint]. 2026. arXiv:2606.02668.
  35. Mazur K. Hack in a box: portable Docker labs for teaching web application security. In: International Conference on Computational Science. Cham: Springer Nature Switzerland; 2026. p.516-30.
  36. Miguel CJ, Izal M. CTF as a service: a reproducible and scalable infrastructure for cybersecurity training [Preprint]. 2026. arXiv:2603.22511.
  37. Canedo A. paper.json: a coordination convention for LLM-agent-actionable papers [Preprint]. 2026. arXiv:2605.16194.
  38. LogJack SH. Indirect prompt injection through cloud logs against LLM debugging agents [Preprint]. 2026. arXiv:2604.15368.
  39. Kozachok AV, Nazimov AM, Magomedov SG. 2DSL: LLM-based code generation for domain-specific languages [Preprint]. 2026. arXiv:2606.22586.
  40. Zhang Z, Xu Y, Liang J, Li W, Chen X, Qian L, et al. RepoZero: can LLMs generate a code repository from scratch? [Preprint]. 2026. arXiv:2605.07122.
  41. Madatha P. A deterministic control plane for LLM coding agents [Preprint]. 2026. arXiv:2606.26924.
  42. Shan Z, Xin J, Zhang Y, Xu M. Don’t let the claw grip your hand: a security analysis and defense framework for OpenClaw [Preprint]. 2026. arXiv:2603.10387.
  43. Xia H, SlotGuard PY. Stop oversharing private local context in LLM agent transcri [Preprint]. 2026. arXiv:2607.17147.
  44. Apostu AM, Preda A, Damir AD, Bolocan D, Ionescu RT, Croitoru I, et al.
  45. Paolillo A. Pythainer: composable and reusable Docker builders and runners for reproducible research. J Open Source Softw. 2026;11(117):9059. doi:10.21105/joss.09059.
  46. Akbar S, El-Hamalawy AF. From Bash to boards: integrating project-based learning in a foundational embedded systems course. 2026 ASEE Annual Conference & Exposition, Charlotte, NC, USA. 2026. doi:10.18260/1-2--59550.
  47. Carrillo B, Mallona I. Denet, a lightweight command-line tool for process monitoring in benchmarking and beyond [Preprint]. 2025. arXiv:2510.13818.
  48. Mudduluru S, Kittur J, Mackay SN. Evaluating the role of research-inspired projects in lower-division programming course. 2026 ASEE Annual Conference & Exposition, Charlotte, NC, USA. 2026. doi:10.18260/1-2--59581.
  49. Chen S, Wang L, Yang X, Liu Z, Cong Y, Ji Y, et al. TUA-Bench: a benchmark for general-purpose terminal-use agents [Preprint]. 2026. arXiv:2606.28480.
  50. Chu Z, Hu J, Jiang X, Zou P, Li H, Peng C, et al. Terminal world: benchmarking agents on real-world terminal tasks [Preprint]. 2026. arXiv:2605.22535.
  51. Meng L, Feng H, Shumailov I, Fernandes E. Cellmate: sandboxing browser AI agents [Preprint]. 2025. arXiv:2512.12594.
  52. Piao Y, Min H, Su H, Zhang L, Wang L, Yin Y, et al. Agent Bay: a hybrid interaction sandbox for seamless human-AI intervention in agentic systems [Preprint]. 2025. arXiv:2512.04367.
  53. Wu Y, Roesner F, Kohno T, Zhang N, Iqbal U. IsolateGPT: an execution isolation architecture for LLM-based agentic systems [Preprint]. 2024. arXiv:2403.04960.
  54. Yan B. Fault-tolerant sandboxing for AI coding agents: a transactional approach to safe autonomous execution [Preprint]. 2025. arXiv:2512.12806.
  55. Dong Y, He J, Liu S, Hou Y, Du D, Xu Z, et al. DeltaBox: scaling stateful AI agents with millisecond-level sandbox checkpoint rollback [Preprint]. 2026. arXiv:2605.22781.
  56. Ruan Y, Dong H, Wang A, Pitis S, Zhou Y, Ba J, et al. Identifying the risks of LM agents with an LM-emulated sandbox [Preprint]. 2023. arXiv:2309.15817. doi:10.48550/arXiv.2309.15817.
  57. Debenedetti E, Zhang J, Balunovic M, Beurer-Kellner L, Fischer M, Tramèr F. AgentDojo: a dynamic environment to evaluate prompt injection attacks and defenses for LLM agents. In: Advances in Neural Information Processing Systems 37. 2024. p.82895-82920. doi:10.52202/079017-2636.
Support