How Secure Is C and C++ Code Generated by Large Language Models?

2 minute read

Published:

Large Language Models (LLMs) are increasingly being used to write, complete, and debug software. These tools can improve developer productivity, but an important question remains: is the generated code secure enough to be used in real applications?

This blog post presents our paper, https://doi.org/10.1145/3748522.3780027, published at the 41st ACM/SIGAPP Symposium on Applied Computing.

Why is this problem important?

Code generated by an LLM may be syntactically correct and may even pass functional tests while still containing security vulnerabilities. This distinction is especially important for C and C++, where unsafe memory operations, incorrect input validation, and weak defensive programming can introduce serious security risks.

Because LLMs learn from large collections of existing source code, they may reproduce insecure programming patterns present in their training material. Developers may also place too much trust in code that appears convincing and is produced almost instantly.

What did we do?

We empirically evaluated C and C++ code generated by ten different language models. The generated programs were analysed using static application security testing tools.

We classified the identified weaknesses using the Common Weakness Enumeration, commonly known as CWE. To understand their potential severity and practical relevance, we also connected the weaknesses with known vulnerabilities represented through Common Vulnerabilities and Exposures, or CVEs.

This approach allowed us to look beyond whether the generated programs worked and examine whether they followed secure programming practices.

What did we find?

Our analysis found a concerning number of weaknesses in the generated code. The results indicate that code generation capability should not be treated as evidence of security.

LLMs can provide useful assistance during software development, but their outputs still require careful review, security analysis, and testing. Developers should treat generated code in much the same way as code obtained from an unknown or untrusted source.

Perspective

The key message from this work is not that developers should avoid using LLMs. Instead, secure use requires appropriate safeguards. Static analysis, dependency checking, compiler protections, code review, and security testing should be integrated into any workflow that uses AI-generated code.

As these tools become more deeply integrated into software development, evaluations should consider security alongside functional correctness and productivity.

The paper was co-authored with Muhammad Usman Shahid and Rajiv Ranjan. The prompts, generated programs, and supporting analysis artefacts are available in the associated research repository.