Table of Contents
    Get answers from the Community
    Join discussions

    “You’re not going to lose your job to an AI, but you’re going to lose your job to someone who uses AI,” Jensen Huang, Founder and CEO at Nvidia.

    Any human work that involves computers is a candidate for automation with AI. Software analysis is no exception, and LLMs can augment software composition analysis (SCA) and open source audits. However, such models neither supplant world-class SCA tools nor human expertise as part of M&A audits in identifying open source risk.

    One of the key reasons to understand the composition of a codebase is to ensure that all open source components in the code are properly and compatibly licensed. Research shows that this is a distinctly weak area for LLMs.

    In LiCoEval: Evaluating LLMs on License Compliance in Code Generation, the authors describe testing 14 LLM models’ abilities to identify the licenses of open source code those models had copied into code they generated. This is a best case scenario for the LLMs—identifying the licensing of code they themselves introduce in software. At best, the models are only 70% to 80% accurate at identifying permissive licenses. Curiously, they are much worse at identifying copyleft-licensed code. In fact, the best of the breed was only 40% accurate with those most problematic licenses.

    Now flip the scenario. LiCoEval tested LLMs on the easiest version of this problem: identifying the license of a bit of code the model itself just wrote, using content it had memorized well enough to reproduce almost verbatim.

    The real-world audit problem is harder in every dimension. You’re asking it to recognize open source licenses and components in a large codebase written by others. If the model can’t reliably get the license right under the best possible conditions, there’s no reason to expect it to do even as well cold, on code it didn’t write and has never seen.

    The copyleft license detection gap LLMs can’t close

    The researchers assume that the gap is due to how the models were trained or perhaps even untrained after the fact. LLMs are weak on licensing and tend to copy/paste third-party code, and coding assistants introduce the risk of improperly licensed snippets. To address this concern with coding assistants, providers are increasingly excluding copyleft code from training data altogether to avoid reproducing it.

    That may be justifiable risk management for the model builder, but out goes the baby with the training bathwater. As a result, such models are not able to identify those components because they were not trained to. Our own empirical experiments suggest 75% to 80% identification accuracy from the best models. From a legal perspective, copyleft-licensed components are the ones that are most important to identify, so this is a huge gap. Put on your CISO hat, and it’s clear that missing components means falling short at identifying vulnerabilities as well.

    Why purpose-built SCA tools beat LLMs at scale

    Further, point an LLM at a real codebase—10 million lines is typical for our audits—and you’re not just getting suboptimal results, you’re burning enormous amounts of compute doing it. A purpose-built matching tool can fingerprint a file against a comprehensive knowledgebase efficiently; asking a language model to reason its way through the same file, repeatedly, across an entire repository, is extremely expensive by comparison. LLMs can complement but not replace SCA tools.

    An alternative to a large general-purpose model is to employ a smaller, specialized one (an SLM) tuned to this narrower task. That can work, but it isn’t a weekend project. It takes a team with real audit domain expertise to know what “right” looks like, and the tooling experience to build and validate a smaller model against it. Building a customized model without that expertise just gets you a cheaper way to be wrong.

    The right way to run an open source audit: SCA + AI + experts

    AI models can certainly support SCA, but with limits.

    Black Duck Audits and Black Duck® SCA are smartly leveraging AI. Using an enormous KnowledgeBase™ curated over two decades, our SCA tools produce the best results in the industry.

    For high-risk scenarios like M&A due diligence, our expert auditors review and curate the results. Complementing that process, an LLM (or SLM) can add observations, identify details a reviewer might otherwise have to dig for, and help our audit team curate results faster. That’s a real productivity gain—a good thing given the way that coding assistants are driving up the size of codebases. So LLMs support the work but don’t eliminate the need for human expert involvement. Lose that, and wrong component identifications quietly and confidently lead to wrong licenses, with nothing flagging that anything is off.

    The most powerful approach to open source analysis is the combination of sophisticated SCA tools with an enormous knowledgebase, complementary AI assistance, and human experts to assure the results.

    To borrow from Jensen Huang: You’re not going to lose your job to an AI, but you might if you lean too heavily on one to identify open source risks in software.