EMPIRICAL ANALYSIS OF PHASE-BASED BEHAVIOR OF AI CODING AGENTS
DOI:
https://doi.org/10.18372/2310-5461.71.21450Keywords:
AI coding agents, large, language models, context engineering, SWE-bench, phase-based analysis, generalized linear model, transition matrices, OpenHands, SWE-agent, AgentlessAbstract
To fix bugs in program code automatically, developers increasingly rely on AI coding agents, including
OpenHands, SWE-agent, and Agentless. These tools take a similar route to a solution, searching for the relevant
fragment in the repository, making changes, and verifying them with tests, yet each agent does this in its own way.
How exactly agents divide effort between these steps, and whether this division relates to the outcome, has not been
studied enough. To find out, 4,300 trajectories of these three architecturally distinct agents were analysed on SWE
bench (Software Engineering Benchmark) Verified. Using tool-call signatures, each agent action was assigned to
one of three phases, namely Localization (finding relevant code), Patching (editing code), or Validation
(running tests). Phase distributions, transitions between phases, and file types were then compared for solved and
unsolved trajectories. A Generalized Linear Model (GLM) showed that the odds of success rise when an agent
validates its changes more (OR = 1.037, p < 0.001) and switches phases more often (OR = 7.843, p < 0.001), and fall
when the number of actions grows too large (OR = 0.984, p < 0.001). The transition matrices revealed two distinct
strategies. In the iterative one the agent tests after almost every edit (OpenHands, 50.7% of post-patch transitions go
to validation), while in the targeted one it makes several edits in a row before checking the result (SWE-agent, 51.4%
repeated edits). Successful trajectories are generally shorter, contain more validation, and switch phases more
often. These findings provide an empirical basis for tailoring agent context to the phase of the task.
References
Yang J., Jimenez C. E., Wettig A., et al. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. Advances in Neural Information Processing Systems (NeurIPS 2024). 2024. URL: https://arxiv.org/ abs/2405.15793.
Wang X., Li B., Song Y., et al. OpenHands: An Open Platform for AI Software Developers as Generalist Agents. International Conference on Learning Representations (ICLR 2025). 2024. URL: https://arxiv.org/abs/2407.16741.
Xia C. S., Deng Y., Dunn S., Zhang L. Agentless: Demystifying LLM-based Software Engine-ering Agents. ACM International Conference on the Foundations of Software Engineering (FSE 2025). 2024. URL: https://arxiv.org/abs/2407. 01489.
Jimenez C. E., Yang J., Wettig A., et al. SWE-bench: Can Language Models Resolve Real-World GitHub Issues? International Conference on Learning Representations (ICLR 2024). 2024. URL: https://arxiv.org/abs/2310.06770.
Du Y., Tian M., Ronanki S., et al. Context Length Alone Hurts LLM Performance Despite Perfect Retrieval. Findings of the Association for Computational Linguistics: EMNLP 2025. 2025. URL: https://arxiv.org/abs/2510.05381.
Gu W., Chen J., Wang Y., et al. What to Retrieve for Effective Retrieval-Augmented Code Generation? An Empirical Study and Beyond. arXiv preprint. 2025. URL: https://arxiv.org/abs/2503.20589.
Gloaguen T., Mündler N., Müller M., Raychev V., Vechev M. Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents? ICLR 2026 Workshop on Memory for LLM-Based Agentic Systems. 2026. URL: https://arxiv.org/abs/2602.11988.
Lindenbauer T., Slinko I., Felder L., et al. The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management. NeurIPS 2025 Workshop on Deep Learning for Code (DL4C). 2025. URL: https://arxiv.org/abs/2508.21433.
Li H., Tang Y., Wang S., Guo W. PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification. International Conference on Machine Learning (ICML 2025). 2025. URL: https://arxiv.org/abs/ 2502.02747.
Mei L., Yao J., Ge Y., et al. A Survey of Context Engineering for Large Language Models. arXiv preprint. 2025. URL: https://arxiv.org/abs/2507. 13334.
Tao Y., Li Y., Qin Y., Liu Y. Retrieval-Augmented Code Generation: A Survey with Focus on Repository-Level Approaches. arXiv preprint. 2025. URL: https://arxiv.org/abs/2510. 04905.
Li H., Zhu L., Zhang B., et al. ContextBench: A Benchmark for Context Retrieval in Coding Agents. arXiv preprint. 2026. URL: https://arxiv.org/ abs/2602.05892.
Zhu J., Wu J., Hu M., et al. SWE Context Bench: A Benchmark for Context Learning in Coding. arXiv preprint. 2026. URL: https://arxiv.org/abs/ 2602.08316.
Liu T., Xu C., McAuley J. RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems. International Conference on Learning Representations (ICLR 2024). 2024. URL: https://arxiv.org/abs/2306.03091.
Blinn A., Li X., Kim J. H., Omar C. Statically Contextualizing Large Language Models with Typed Holes. Proceedings of the ACM on Programming Languages (OOPSLA 2024). 2024. Vol. 8. P. 468–498. URL: https://doi.org/ 10.1145/3689728.
Liu S., Chen Y., Krishna R., et al. Process-Centric Analysis of Agentic Software Systems. arXiv preprint. 2025. URL: https://arxiv.org/abs/ 2512.02393.
Yang S., Li Y., He S., et al. Phase-Aware Mixture of Experts for Agentic Reinforcement Learning. arXiv preprint. 2026. URL: https://arxiv.org/abs/2602.17038.
Liu S., Yang J., Jiang B., et al. Context as a Tool: Context Management for Long-Horizon SWE-Agents. arXiv preprint. 2025. URL: https://arxiv.org/abs/2512.22087.
Wang Y., Shi Y., Yang M., et al. SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents. arXiv preprint. 2026. URL: https://arxiv.org/abs/2601.16746.
Bui N. D. Q. Building Effective AI Coding Agents for the Terminal: Scaffolding, Harness, Context Engineering, and Lessons Learned. arXiv preprint. 2026. URL: https://arxiv.org/ abs/2603.05344.
Nebius Team. Nebius SWE-agent Trajectories Dataset. Hugging Face Hub. 2025. URL: https://huggingface.co/datasets/nebius/SWE-agent-trajectories.
Nebius Team. Nebius OpenHands Trajectories Dataset. Hugging Face Hub. 2025. URL: https://huggingface.co/datasets/nebius/SWE-rebench-openhands-trajectories.
McCullagh P., Nelder J. A. Generalized Linear Models. 2nd ed. London: Chapman and Hall, 1989.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Олександр Дроздюк

This work is licensed under a Creative Commons Attribution 4.0 International License.
The scientific journal adheres to the principles of Open Access and provides free, immediate, and permanent access to all published materials without financial, technical, or legal barriers for readers.
All articles are published in Open Access under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
Copyright
Authors who publish their works in the journal:
-
retain the copyright to their publications;
-
grant the journal the right of first publication of the article;
-
agree to the distribution of their materials under the CC BY 4.0 license;
-
have the right to reuse, archive, and distribute their works (including in institutional and subject repositories), provided that proper reference is made to the original publication in the journal.



