Architectures of Error: A Philosophical Inquiry into Human and AI Code

· SpringerLink

9 min read Original article ↗

Abstract

With the rise of generative AI (GenAI), Large Language Models are increasingly employed for code generation, becoming active co-authors alongside human programmers. Focusing specifically on this application domain, this paper articulates distinct “Architectures of Error” to ground an epistemic distinction between human and artificial code generation. Examined through their shared vulnerability to error, this distinction reveals fundamentally different causal origins: human-cognitive versus artificial-stochastic. To develop this framework and substantiate the distinction, the analysis draws critically upon Dennett’s mechanistic functionalism and Rescher’s methodological pragmatism. I argue that a systematic differentiation of these error profiles raises critical philosophical questions concerning semantic coherence, security robustness, epistemic limits, and control mechanisms in human-AI collaborative software development. The paper also utilizes Floridi’s Levels of Abstraction to provide a nuanced understanding of how these error dimensions interact and may evolve with technological advancements. This analysis aims to offer philosophers a structured framework for understanding code generation in the context of GenAI’s epistemological challenges, shaped by its architectural foundations, while also providing software engineers with a basis for more critically informed engagement.

Access this article

Log in via an institution

Subscribe and save

  • Starting from 10 chapters or articles per month
  • Access and download chapters and articles from more than 300k books and 2,500 journals
  • Cancel anytime
View plans

Buy Now

Price excludes VAT (USA)
Tax calculation will be finalised during checkout.

Instant access to the full article PDF.

Fig. 1

Similar content being viewed by others

Data Availability

No data or supplementary materials are available.

Notes

  1. The philosophical implications of LLMs functioning primarily as pattern detectors—a capability integral to, but not exhaustive of, reasoning—remain understudied, despite their demonstrated competence.

  2. A detailed prompt and concurrent test harness (Appendix A) reveal critical failures even in state-of-the-art model outputs.

  3. Temperature is a parameter controlling randomness in LLM output; lower values yield more focused, less varied (though often still non-deterministic) responses.

  4. An area that inherently engages with hybridization is Metaheuristics. Rather than a single algorithm, a metaheuristic is a general framework for working with various heuristic methods, especially in the context of optimization problems. These approaches can often be combined with exact methods—procedures that guarantee finding the optimal solution for a given problem instance, such as Integer Linear Programming (Blum & Raidl, 2018).

    Their epistemic challenges echo those of GenAI: How can algorithms that rely on random operators achieve strong empirical performance without a theoretical account of their behavior?

References

  • Abbassi, A. A., Silva, L. D., Nikanjam, A., & Khomh, F. (2025). Unveiling inefficiencies in llm-generated code: Toward a comprehensive taxonomy. Retrieved from arxiv:2503.06327

  • Baquero, C. (2025). The last solo programmers. https://cacm.acm.org/blogcacm/ the-last-solo-programmers/. Accessed 26 Apr 2025

  • Beschastnikh, I., Wang, P., Brun, Y., & Ernst, M. D. (2016). Debugging distributed systems: Challenges and options for validation and debugging. Queue, 14(2), 91–110. Retrieved from https://doi.org/10.1145/2927299.2940294

  • Bianchini, F. (2025). Generative artificial intelligence: A concept in progress. Philosophy & Technology, 38(2), 46. Retrieved from https://doi.org/10.1007/s13347-025-00875-8

  • Blum, C., & Raidl, G. R. (2018). Hybrid metaheuristics: Powerful tools for optimization (1st ed.). Incorporated: Springer Publishing Company.

    Google Scholar 

  • Caporuscio, C. (2021). Introspection and belief: Failures of introspective belief formation. Review of Philosophy and Psychology. Retrieved from https://doi.org/10.1007/s13164-021-00585-y

  • Dennett, D. (1971). Intentional systems. Journal of Philosophy, 68(February), 87–106. https://doi.org/10.2307/2025382

    Article  Google Scholar 

  • Dennett, D. (2017). From bacteria to bach and back: The evolution of minds.

  • Floridi, L. (2008). The method of levels of abstraction. Minds and Machines, 18(3), 303–329. Retrieved from https://doi.org/10.1007/s11023-008-9113-7

  • Floridi, L. (2019). What the near future of artificial intelligence could be. Philosophy & Technology, 32(1), 1–15. Retrieved from https://doi.org/10.1007/s13347-019-00345-y

  • Floridi, L. (2023). Ai as agency without intelligence: on chatgpt, large language models, and other generative models. Philosophy & Technology, 36(1), 15. Retrieved from https://doi.org/10.1007/s13347-023-00621-y

  • Günther, M., & Kasirzadeh, A. (2022). Algorithmic and human decision making: for a double standard of transparency. AI & Society, 37(1), 375–381. Retrieved from https://doi.org/10.1007/s00146-021-01200-5

  • Hähnel, M., & Hauswald, R. (2025). Trust and opacity in artificial intelligence: Mapping the discourse. Philosophy & Technology, 38(3), 115. Retrieved from https://doi.org/10.1007/s13347-025-00947-9

  • Hosseini, P., Castro, I., Ghinassi, I., & Purver, M. (2024). Efficient solutions for an intriguing failure of llms: Long context window does not mean llms can analyze long sequences flawlessly. Retrieved from arxiv:2408.01866

  • Huang, D., Xie, X., Zhang, J., Chen, J., Bu, Q., & Cui, H. (2024). Bias testing and mitigation in llm-based code generation. Retrieved from arxiv:2309.14345

  • Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., & Liu, T. (2025). A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems, 43(2), 1–55. Retrieved from https://doi.org/10.1145/3703155

  • Huynh, N., & Lin, B. (2025). Large language models for code generation: A comprehensive survey of challenges, techniques, evaluation, and applications. Retrieved from arxiv:2503.01245

  • Kosinski, M. (2024). Evaluating large language models in theory of mind tasks. Proceedings of the National Academy of Sciences, 121(45), e2405460121. https://doi.org/10.1073/pnas.2405460121. Retrieved from https://www.pnas.org/doi/abs/10.1073/pnas.2405460121

  • McKendrick, J. (2025). Will AI replace software engineers? It depends on who you ask - zdnet.com. https://www.zdnet.com/article/will-ai-replace-software-engineers -it-depends-on-who-you-ask/. Accessed 16 May 2025

  • Novikov, A., Vũ, N., Eisenberger, M., Dupont, E., Huang, P. -S., Wagner, A. Z., ..., Deepmind, G. (2025). AlphaEvolve : A coding agent for scientific and algorithmic discovery.

  • Ouyang, S., Zhang, J.M., Harman, M., & Wang, M. (2025). An empirical study of the non-determinism of chatgpt in code generation. ACM Transactions on Software Engineering and Methodology, 34(2). Retrieved from https://doi.org/10.1145/3697010

  • Pierce, B. C. (2002). Types and programming languages (1st ed.). The MIT Press.

  • Rescher, N. (1992). Rationality: A philosophical inquiry into the nature and the rationale of reason the clarendon library of logic and philosophy. Philosophy and Rhetoric, 25(1), 82–84.

    Google Scholar 

  • Rescher, N. (2003). Epistemology: An introduction to the theory of knowledge. New York: State University of New York Press.

    Book  Google Scholar 

  • Rescher, N. (2017). Value reasoning: On the pragmatic rationality of evaluation. Berlin: Springer International Publishing.

    Book  Google Scholar 

  • Robeyns, M., Szummer, M., & Aitchison, L. (2025). A self-improving coding agent. Retrieved from arxiv:2504.15228

  • Shanahan, M. (2024). Talking about large language models. Commun. ACM, 67(2), 68–79. Retrieved from https://doi.org/10.1145/3624724

  • Shapira, N., Levy, M., Alavi, S. H., Zhou, X., Choi, Y., Goldberg, Y., & Shwartz, V. (2024). Clever hans or neural theory of mind? stress testing social reasoning in large language models. In Y. Graham, & M. Purver (Eds.), Proceedings of the 18th conference of the european chapter of the association for computational linguistics (volume 1: Long papers) (pp. 2257–2273). St. Julian’s, Malta: Association for Computational Linguistics. Retrieved from https://aclanthology.org/2024.eacl-long.138/

  • Simon, J. (2015). Distributed epistemic responsibility in a hyperconnected era. In L. Floridi (Ed.), The onlife manifesto: Being human in a hyperconnected era (pp. 145–159). Cham: Springer International Publishing. Retrieved from https://doi.org/10.1007/978-3-319-04093-6_17

  • Thompson, K. (1984). Reflections on trusting trust. Commun. ACM, 27(8), 761–763. Retrieved from https://doi.org/10.1145/358198.358210

  • Zerilli, J., Knott, A., Maclaurin, J., & Gavaghan, C. (2019). Transparency in algorithmic and human decision-making: Is there a double standard? Philosophy & Technology,32(4), 661–683. Retrieved from https://doi.org/10.1007/s13347-018-0330-6

Download references

Acknowledgements

I am grateful to Raymond Turner and William J. Rapaport, whose books helped me see that computer science can–and should–be understood as more than a purely technical discipline.

The research presented in this paper was part of the R&D project PID2022-138283NBI00, funded by MICIU/AEI/10.13039/501100011033 and “FEDER – A way of making Europe” (Camilo Chacón Sartori).

Funding

R&D Project PID2022-138283NBI00, funded by MICIU/AEI/10.13039/501100011033 and “FEDER – A way of making Europe” (Camilo Chacón Sartori).

Author information

Authors and Affiliations

  1. Artificial Intelligence Research Institute (IIIA-CSIC), Bellaterra, 08193, Barcelona, Spain

    Camilo Chacón Sartori

  2. Institut Catalá de Nanociéncia i Nanotecnologia (ICN2), Bellaterra, 08193, Barcelona, Spain

    Camilo Chacón Sartori

Authors

  1. Camilo Chacón Sartori

Contributions

Not applicable

Corresponding author

Correspondence to Camilo Chacón Sartori.

Ethics declarations

Ethics approval and consent to participate

Not applicable

Competing interests

The author declares that they have no competing interests.

Appendix A A Prompt that Stresses All Facets of “Architectures of Error”

Appendix A A Prompt that Stresses All Facets of “Architectures of Error”

Author’s note: Paradoxically, I structured this prompt with the help of an LLM; however, the fact that an LLM can craft a complex prompt does not mean it is equally competent at solving it.

The following prompt poses a challenge even for programmer experts, as it brings together multiple software engineering concepts—such as invariants, concurrency, deep abstraction, fine-grained synchronization, multiple interacting algorithms, and semantic coherence—that are notoriously difficult to reason about when entangled. It demands the design of a complex internal architecture rather than the mere implementation of a function. Even today, such scenarios remain non-trivial for the most advanced generative models (Gemini-2.5-Pro, GPT-4o, Llama-4, Claude-4).

As prompt complexity grows, models may handle parts correctly, but global inference errors become more likely. Real-world software often involves even messier cases—integrating new code with legacy systems and entangled concepts.

figure a
figure b
figure c
figure d
figure e
figure f

I consider a crucial aspect of the prompt design to be the explicit definition of the class initialization and core method signatures. Without such clear specifications, a generative model is more prone to produce irrelevant outputs, as it navigates a significantly larger search space to satisfy the request. Analogous to instructing humans, constraining the output of a GenAI model necessitates clear and well-defined instructions.

Nevertheless, increasing the number of distinct concepts or complex constraints within a single prompt generally elevates the probability of model failure. Consequently, decomposing a complex request into more specific, modular prompts is often a more effective strategy for guiding GenAI. However, even this level of modularization falls short of fully addressing an LLM’s inherent struggle with the intricate logic and hierarchical complexity posed by challenges like the AdaptiveHierarchicalTaskAssigner. A similar situation arises with human programmers, who, depending on their experience, may only be able to tackle certain parts of the problem effectively.

Yet the question remains: how far can we constrain a prompt to fulfill a requirement without making the generated output invalid or ineffective? As Aristotle might say when discussing moral virtue in humans, when we create instructions for a GenAI model to generate code, We must ask: where lies the mesotes—the virtuous middle ground—between being explicit and giving the model greater freedom to infer code?

About this article

Check for updates. Verify currency and authenticity via CrossMark

Cite this article

Sartori, C.C. Architectures of Error: A Philosophical Inquiry into Human and AI Code. Philos. Technol. 39, 55 (2026). https://doi.org/10.1007/s13347-026-01056-x

Download citation

  • Received:

  • Accepted:

  • Published:

  • Version of record:

  • DOI: https://doi.org/10.1007/s13347-026-01056-x

Keywords