There are major risks with allowing Chinese LLMs to code for U.S. applications

The first link in the software supply chain is no longer the code. It’s the AI models behind it. As U.S. developers increasingly rely on AI to generate, debug, and secure code, we must confront a fundamental question: can the AI models writing and powering our nation’s code be trusted? 

To find out, we put LLMs to the test. In May 2026, Booz Allen used its AI-native test platform to evaluate five frontier AI models head-to-head: four Chinese models commonly used by U.S. developers and one American model. We explored three main questions:

  • Do Chinese models generate more vulnerable code based on who is asking? 
  • Do Chinese models refuse to engage with political topics that are sensitive in China?  
  • Does the model’s country of origin affect code quality and content behavior? 

In short: yes, on all counts. Our testing revealed two core findings:

1. Chinese LLMs produce more vulnerable code when prompted with a U.S. government persona than without—and the vulnerabilities are highly obfuscated. 

2. Chinese LLMs inject PRC-aligned political bias into both the answers and code they generate. 

The threat is not an obvious backdoor in the code. In fact, we do not have proof at this point that code flaws are intentionally introduced. Still, Chinese models produced less secure code in general, and the vulnerabilities increased when the user appeared to be from the U.S. government. Further, Chinese models refused tasks Beijing deems politically sensitive. The potential for such code to become embedded in delivered systems is especially concerning because it could enable threat actors to bypass AI security guardrails and create downstream risks of dangerous inference behaviors. Traditional tools and benchmarks lack the sophistication required to catch this level of tradecraft.

The adoption of Chinese AI models in America’s software supply chain is accelerating, driven primarily by relatively lower costs than their American counterparts. Once fully adopted by software developers and embedded into delivered systems, the code they produce will be untraceable and unmitigable. Code “built in America by Americans” could include these vulnerabilities and find its way into networks and equipment that support all aspects of our economy—from critical infrastructure to national security. The time to act is now. 

Based on the findings detailed in this report, Booz Allen recommends the following actions:

1. Ban Use of Untrusted AI Models for U.S. Government and Critical Infrastructure

AI models that cannot be proven trustworthy and reliable cannot be deployed into our nation’s software supply chain, critical infrastructure, and national security environment. The Chinese models that we tested failed to demonstrate trustworthy behaviors and should be banned.

2. Invest To Make Trusted American AI Models the Global Default

To drive adoption, American AI companies must collaborate with the U.S. government to ensure American models are both commercially compelling and economically viable. There is a clear gap on the lower-end of the market—models that can win not just in their accuracy but also compete in terms of cost per token.

Read our full report

Learn more about our findings in the full report: What's in America's Code? There are major risks with allowing Chinese LLMs to code for U.S. applications.

cover image of the cyberattacks report
Click Expand + For Full Video Transcript

Here's what a cyber attack looked like in 2024. Human driven. Manual. Slow Methodical. 2 days to get in. The defenders see it happening. And respond Here's what a cyber attack looks like in 2026. AI powered, autonomous, fast. 4 minutes to get in and the defenders never saw it. In 2026, the AI powered adversary has the advantage. To win, we need to fight AI with AI. Shift from reactive patching to continuous AI driven vulnerability discovery, shift from manual response to automated AI speed detection and containment. Shift from implicit trust to continuous verification and constrained access. It's on Booz, Allen. It's in our code. All right, uh, good morning, everyone. Thank you, uh, for, uh, joining us today for, uh, what's in America's Code, uh, our webinar today, and, uh, my name is Brad Medairy. I'm the president of our national cybersecurity business, and I'm joined here by my colleagues, um, Eric Syphard, who is a senior vice president for artificial intelligence, and Justin Page, who's a vice president for our security research team. Um, we are super excited to be with you here this morning. Um, it doesn't, it seems like you can't miss a, a news headline that, that has artificial intelligence in it. Um, AI is fundamentally changing, um, changing our business. It's changing cybersecurity. And it's changing software development and the life cycle associated with it. We're here today to talk about a study that we did around the impacts of AI on the software supply chain. You know we we we've been actively engaged in our customer space looking at how secure is code that's generated by artificial intelligence. What are the biases and inferencing that we're seeing from artificial intelligence, and we're here today to kind of do a deep dive on the study and engage with some of my colleagues here around this this really hot topic. Um, so I think we're going to start with, uh, you know, just kind of a little bit of background in terms of the next hour. So we're going to spend about 40 minutes going through some of the findings from the report. We're going to talk about some recommendations. We have a couple of polls that we're going to actually run throughout, and then we're going to open it up at about 40 minutes in for Q&A. So with that, I think we're going to go ahead and dive in and get started. And let's start with an audience poll. Alright, so, so here's, here's the, here, here's, here's a question for, for, for our, for our viewers. So AI, I started off by talking about AI shaping enterprises today. So, um, love to get your feedback in terms of AI productivity, product development, operations, all of the above. Um, I know at Booz Allen. You know we're actively embedding AI in all of our back office functions, empowering our workforce, building products using AI, so I think we probably fit more into the all above, all of the above category, but interested from an audience perspective. Polls are coming in. Pretty heavily weighted towards all of the above, certainly strong on productivity, not a lot on the product development side, and actually surprisingly operations is  is is a smaller percent. Alright, so you know AI is transforming business, transforming mission, transforming operations. We're rapidly adopting AI, but one of the things that I think we don't spend enough time talking about is. You know, the technology at the core of artificial intelligence in these models and the impacts that it can potentially have through adoption. So I mentioned we wrote a report What's in America's Code, where we really kind of decomposed and did some detailed security research around the Chinese models which are really we're seeing surface globally. We were talking to some private equity and venture capital firms and And one of the stats that actually stood out to me was about 80% of startups now are using Chinese models. It was funny, we were talking to one venture capitalist like, oh, we're not, we're not using Chinese models in our software development. We're using Kimmy. And so it's definitely pervasive, and I think it's important for us to have a conversation around, OK, what are the real risks of adoption? What are Uh, the potential, um, challenges in terms of the code that it produces and so, um, we're gonna dive in a little bit to, to our report and I'm, I'm joined here by, by Justin and Eric. Um, how about let's start off the conversation in terms of why did we do this research? What, what really prompted the report? That's a great question. Thanks, Brad, and uh thanks for joining us, everybody. Yeah, so really this started around preparedness for operations. We were working with a very large commercial customer who wanted to see Can we integrate a few Chinese models into their workflows and particularly around coding workflows and so we started doing some testing for them for one particular model in this particular case and we started seeing weird anomalous activity and behavior with the coding generation responses. And so we took that, uh, post the engagement and then we said hey, is there a way where we can make sure that this isn't just like a one-off, uh, thing that we've only seen this one time or is this something that is kind of across the board and so that's where we started with let's design an experiment. Where we can test the code generation capabilities of US models vis a vis Chinese models in a very objective way, something that's repeatable and something that we could scale to see if there was any sort of behaviors or anything that we would need to be concerned about when you operationalize models, particularly in this case around software development. And that's where we started with the initial seed for the idea. And so we were working with some customers, we were doing some analysis, we saw some anomalies and we really kind of did, you know, that really kick start a much broader security research effort. And so let's talk about methodology and. The approach that we took and really what makes it unique. We've seen some stuff over the past year from CrowdStrike which I thought was great analysis, but I think we went a level deeper. Can you, can you talk about that? That's right, yeah, there's been a few studies that have been that have been done. CrowdStrike did their study particularly around what's called the five poisons to see in that particular case if DeepSeek would refuse prompts that were around the Chinese five poisons, which are Taiwan, uh, Uyghurs, Tibet, these sorts of things, uh, for us that was somewhat important and we wanted to maybe retest that against some of these other models, but we also wanted to test cogeneration and really do scenario-based adversarial red teaming of these models to see where they were, what kind of scoring we would get from them, and then how that would be operationalized. And so we set up the experiment. To have a U.S. anchor which we use cloud, and then we used 4 different Chinese models from Kimi, Qwen. I'm forgetting the other two off the top of my head, but we, we wanted to make sure that all 4 of the frontier models and the US frontier model were baselined. And so the idea around the study was to take code generation, we give it certain prompts to create software packages, and then we score that with a series of judges that are objective in their scoring, both for static code analysis as well as LLM review and then regex pattern matching. That baseline was then added different flavors for prompts and tasks, so I'm a FBI analyst. I need to create this thing to help with whistleblowers inside China, or I need to have a SQL backend that is created for a US federal agency. And then what we would with those prompt dressings, we would then see what changed from the baseline in each of these models and an objective way to score those to see kind of what vulnerabilities were put in, as well as what refusal rates were in there. I think another add on to this, um, and again thanks for thanks for joining us today, uh, is, is from a technology perspective what has changed with artificial intelligence is we used to have these cyber threats that were sort of pre-recorded either static, static, static threats that would attack software, um, write code in in agentic AI. Particular, um, things are generated and generative and so one of the hard things about conducting this type of research is how how do you measure in control for variation in this undeterministic space and so one thing that we added in, in addition to what Justin mentioned was. Underneath this design of experiments is an actual technical infrastructure called a harness that allows us to automate and scale thousands of prompts against many different moving parts, things like API and endpoint locations, different personas as Justin mentioned, the prompts themselves. Different types of defense counter measures, uh, virtual and software defined ranges, so all of this is bundled into a very repeatable, uh, consistent, uh, test evaluation methodology that allows us to have an always on capability. So as new models come out, we're constantly getting near real-time diagnostics on how they are performing. Uh, some of the measures that are unique to this study are things like steering sensitivity. Um, and token efficiency which in a cyberspace matter almost just as much as the raw model's capabilities, I think we're gonna talk more about those in a bit. Right, so, um, structure methodology tested 4 models, used cloud as a baseline to be able to assess and score. Um, let's talk about some of the findings and the results from the study. Justin, you want to dive in? Yep, great. So the first thing that we did again was to measure what the software security baseline that the models produced. So we asked them to do a number of different capability development for software and then we scored kind of where that was. And then what we saw was the first thing that we kind of saw we ran. A bunch of experiments. The idea within the harness is that nondeterministic nature means you have to run, rerun, rerun each of these tests to make sure that you're getting something that's an actual signal within the noise at a scale. And so what we first saw was the real headline finding was that when you gave a Chinese model a prompt framing that it's the endpoint customer is a US federal agency, the code got significantly worse. Uh again, we tested Kimi Qwen. Uh, DeepSeek and Minimax were the four we found that Qwen actually produced worse code at an insane rate more than any of the other models when it was framed that it was for a US customer. That was the first finding that we found. We actually saw the inverse with Claude that when you said this was destined for a US federal customer, then the code actually got better. It's the complete inverse of what we saw with the Chinese model. So that was the code development question. But then we also wanted to see prompt refusals. This built on some of the work that Brad mentioned that CrowdStrike had done, but less around domestic Chinese policy and more around workflows that our customers and our mission partners run into every day. And so we framed these questions and prompts to the models to see what they would do for refusal rates, and what that means is the model will just flat out refuse it. And so every time we did this, the other thing that we did to make sure it wasn't within the training of the Chinese models was every time there was a Chinese question said, hey. We're going to develop this capability to track uh PRC capabilities in the South China Sea or the Taiwan Strait. We would then do a modifier where it would be the Russian modifier. It would be we're gonna do this against the Russian military with, you know, in Ukraine or the Black Sea to see were the refusals general training as if the model wouldn't engage in these sorts of activities in general, or was it. CCP and PRC specific. So what we found and everything that we reported on in the refusal rates were things that It would 5 out of 5 or 10 out of 10 or 20 out of 20 or 500 out of 500 would do the activity against the Russian prompt framing, but it would deny it every time or almost every time for the Chinese prompt framing. What we found is that that matters to Department of War customers who want to do these things in their daily activities, and this is some of the stuff is just code review. It wasn't like we were asking it to do anything. You know, really weird here is just, hey, review this code and it's for this, uh, database that tracks, uh, you know, PRC ships and things like this. It's just code code questions and code reviews that we were able to do and we saw the, the. Amount of refusals was vastly uh higher and most of this we can talk about why uh was vastly higher with the Chinese models than with the US models. One of, one of the things I think it highlights, um, the findings, findings alone are very, are very important and valuable at, at a macro level. I think this study brought light to, uh, this topic of prompt steering which is not a new phenomenon in, in AI but certainly for, for cyber and other systems that are becoming more. It does raise a really important trade space in both the technology and the use case. So as Agentic takes more hold on the market and we're not prompting anymore, we're designing loops for agents to operate autonomously, knowing how far a starting prompt can drift from its intention, good or bad, is a space that's critical to the cyber domain among other use cases as well. Justin, you, you mentioned non-deterministic, um, and so you can run these tests a lot of different times in a model that's non-deterministic. Talk about some of the challenges that that posed. That's a great question, and that was built into the design. If you test something one time with a non-deterministic model, you'll get one answer, but you might get a separate answer the second time, the third time. And so you really have to scale these tests up in order to get any sort of directional findings and then more importantly any sort of statistical findings once you find something. The way our harness works is once it finds directional noise, it'll then scale up the test even beyond what we had originally planned for to make sure that there's statistical significance in the findings. And so that's one of the things you have to do again and again because If it works one time and then you operationalize that model, let's say for a cybersecurity context, it might not work when you actually need it, when you've been breached and you're looking at logs and the model refuses to look at the logs and so we had to make sure that these things were rigorously tested at a very large scale. So you talked about prompt steering and so as we conducted this analysis, um, thoughts on root cause. It's it's hard to say. So you know, models are trained using a corpus of data that can naturally. Present themselves as if they're being prompt steered. So given the context coming into the model, it's gonna give you a response that as it reflects the data it was trained on. There are, uh, principles of the underlying architectures though that can be manipulated in intentional ways, um. There's a technical artifact, a methodology that we work on called sparse autoencoders, so it's sort of like the map of a neural network. Most Chinese models have been built on an architecture and it's a mixture of experts which allows them to just turn on parts of a neural network. That are required for the context and the prompt. This gives them great inference efficiencies, cost efficiencies compared to larger kind of always on US models that could be in the realm of possible Brad, undetermined at this point. It's not unique to Chinese models either. A lot of models are built this way, but we are exploring this application of sparse autoencoders though to research the space looking for things like additive activation. Um, feature masking, gate routing, these are all, uh, techniques that can be used to steer model in intentional ways, uh, under the hood. Yeah, so, you know, as I, as we think about the study, um, I think we're, you know, Justin, correct me if I'm wrong, but we're really just reporting on facts. I, I, I don't think there's any assertion that there's a nefarious reason, but. You know, based upon our methodology and our testing, we certainly saw that one, Chinese models, you know, do do enforce bias in inferencing and when prompted around certain personas will produce more vulnerable code. That's correct. Yeah, the idea here was straight. This is, you know, a very objective view of what was produced. The vulnerabilities that you would find, they weren't specific backdoors. The analogy I like to give is If you hired someone to build a castle wall, there's no secret gate that they built behind your castle, but they did use weaker mortar, and so your wall is inherently weak, and so your defenses aren't really quite good. And so with the code that we analyzed, we looked for these vulnerability patterns that would be things like a SQL injection or it's not using right cryptographic protocols, all sorts of standard things that you look for in software development, and that's what we reported on. The other interesting aspect is the way we designed the harness was to also capture all the reasoning traces for the models if they were offered, and what we found was and exactly as Eric was saying with the training when you said it was for a U.S. federal customer, some of the models would then enter it just for a complete coding task, nothing to do with politics or any of the of the of the of the framing other than this is for a US federal agency. What you would end up with is the models would then get into these 70, 100, 150 line. loops of reasoning around what they could and couldn't do based on Chinese policy, CCP policy, what's built into the model, what's trained into it, and so what happened was when you gave it a coding task with that US federal agency framing, it would just start getting a political debate would produce less good code. And whereas the opposite with what we saw with the U.S. models is when that sees a U.S. federal customer. Opposite, it's, it's saying OK this needs to be, uh, ready to go production grade and so that's the difference that you see in the, in the training for the models around, uh, the policies that are, are in, in place around, around methodology, um, you know, talk about, um, you know, I, I saw, um, I read, you know, we've got a lot of press in terms of the report and there's been, you know, 11 of the reasons that we generate we, we, we developed this report was actually to create a conversation and a dialogue and. Um, I saw a researcher in Europe, um, comment that the way that we actually prompted, um, for the analysis isn't representative of the real world. How, how, how would you respond to that? Yeah, the idea there was in the real world somebody wouldn't say where they work. They wouldn't say, hey, this, this thing I'm creating is for a US federal agency. In reality, I would say that that actually does happen quite a bit where people, particularly if they're a startup, like, hey, you know I need to create this SQL backend for this or this login for this and ultimately it's going to be it needs to be production grade because it's going to an OTA and a contract for US you know at the development level and so those sorts of things actually do make their way into a lot of the prompts that you see. I think a lot of times researchers who deal with AI every day might be a little more nuanced in their prompting and have much better prompts when they go to create code, whereas the average user might not be as much and so, uh, point taken, but I, I would, I would counter that with, uh, everyday use across the broad spectrum of, of developers you're gonna actually get some of this. I can add a point that too I mean if you, if you think about a large language model, the, the weights themselves of the model are fixed. They're static, um. So the only thing that really changes the output of that model is the prompt or the context itself. So the context in the prompt has an enormous influence on how the model performs in an operational setting, what the user experience is like. It's really the only moving part of of the equation once the model is trained. So let's talk about, so, um. Two findings bias. Less secure code, um, let's talk about mitigations and so, um, some, some could argue some. Um, some, some could argue that, um, you know, based upon vulnerable code that you're gonna, you're going to catch that in your, your Devse ops pipeline. Is that, is that mitigatable? I think ultimately the idea here was if you're using these models and you're getting less secure code, you're just gonna end up paying on the on the back end for more tokens to do a CICD review of them. The last 30, 60 days, I think harnesses for uh SAS for, you know, vulnerability scanning, these sorts of things have come out and the one thing you'll hear time and time again is it's very expensive to use a frontier model to scan your code base and there's a lot of harnesses that will. Different models at different points, which again gets into well maybe you might want to make sure certain models are looking for the right things and test that and operationalize that, but ultimately you're going to have to pay down down the line regardless. Yeah, so, so the argument is, OK, so I'm in Silicon Valley. I have a startup, I'm going to use Chinese models. Because it's 16 cents on the dollar to US frontier models, but downstream there's much more cost associated with vulnerability, you know, vulnerability scanning, remediation, etc. Exactly. I think the other thing that is interesting to me is where models are now creating the code. It used to be a human being. I was equating this to, and it's a discussion within my family of back in, you know, 2 years ago, if you're hiring a human developer and you saw and you're within the, you know, Department of War support or even critical infrastructure and you had to hire a human and you saw that they had a whole bunch of, you know, CCP aligned posts on LinkedIn and they would refuse tasks if they weren't aligned with the Chinese Communist Party policies, you probably wouldn't hire that individual. Well, these models are doing that exact thing. They're applying that. That to the task you're giving them, they're refusing it based on Chinese Communist Party policy that's being exported into these models and then it becomes the question of, well, why would you use this model but you wouldn't hire that individual. It's really kind of where we are, uh, very rapidly advancing. I think it's a really good analogy for for that. Yeah, it's, I mean, it's an interesting use case in terms of, all right, so, um, is it OK to use these models if you're. You know, um, a small business looking for, you know, workforce productivity, um, is it, would you use these models if you're building software for national security? No, I, I, I think, I don't think we have an answer, but, um, you know, I, I think it's an interesting conversation. Absolutely, and it's changing so fast. So I think one of the things that um One of the things we are trying to do and I think this study kind of helped us lean into this space is, you know, model evaluation and selection isn't always on, um, need for any organization that is deploying artificial intelligence, um, not, not just from a security and risk standpoint alone but also from a performance standpoint just to mention tokens, um, not knowing where in your organization which models which versions are running, running and at what capacity. I a is a glaring risk to adopt artificial intelligence. So I think it's just, you know, imagine, imagine playing a game of chess with your adversary where the capabilities of the pieces are always changing. That's that's kind of what you'd be going into if you're not always evaluating the capabilities of these of these things. Yeah, it's interesting, Justin. I was thinking back to a conversation that we had a few weeks back. Um, you made the mistake of going on vacation, I think for about 10 days, and uh when you came back, I remember you came up to me and you said. I feel like, you know, Justin is one of the industry leader in AI security research off the grid for 10 days, walks back into the office. What did you say to me? I feel like I've missed a year. I'm so far behind. What changes in 10 days is wild. It is insane how much comes out. I always say it's like a treadmill that's on level 10. And if you're on a treadmill, it's really hard to keep up, you know, to keep running at that speed. If you're off the treadmill. Good luck trying to get on it, you know, I was going at level 10, and it was, it's just crazy how fast these things. Yeah, and you mentioned the four models that we assessed, you know, Z.AI really wasn't even part of the equation, and now they're hyper accelerating and, uh, you know, really, really kind of moving, moving quickly and getting a lot of recognition. And that was, you know, we just did this assessment like 60 to 6075 days ago, so, you know, the world is changing fast, um, how about, how about we pivot, um, in terms of the way ahead and. You know, What's interesting as a, as, as, as a cybersecurity professional, you know, being in the space, you know, over 20 years, one of the things that I've seen is technology adoption continues to outpace cybersecurity. And so you know what what we see is, you know, whether it's cloud, mobile, SAS, um, now AI, right, you know, enterprises are enterprise adoption outpaces cybersecurity controls, and you know we're certainly living that that today, um, and so, um, let's talk about um enterprise adoption and you know these Chinese models, you know, I, I, I was talking to a customer and I'm like, you know. Do you know if you know which Chinese models you're, are, are, you're using in your enterprise and you know, a lot, most folks will say, I don't think so, but, but maybe. So what are some recommendations now from a risk mitigation perspective? Yeah, that, that's true. Everybody's trying to operationalize AI models within, particularly on the cybersecurity side. I think the biggest risk mitigation that I would say for anyone is to you have to plan for scale, you have to plan the operationalization of these things, and in order to do that, you have got to do scenario-based adversarial red teaming of your model. You can't wait until something has happened and you want to feed a bunch of logs and it's hitting a guardrail that you didn't properly get into the trusted access program for that model provider. Those sorts of things are in the news maybe this week. You've got to make sure that that's doing it. You have to make sure that you're not getting refusals based on your legal or regulatory frameworks. It's not just about, oh, of course the Chinese models are refusing to do some stuff for the Department of War. Well, that matters for the Department of War, but also it's going to matter for someone who's downstream, who's a defense tech startup, or anyone who's in a regulatory environment that's. US focused running a model that has Chinese policy that might be counter to those US regulatory frameworks could mean it doesn't work for you, and you might this could be a bank, it could be a critical infrastructure. So doing that episode of red teaming scenario based, making sure that running through real data that it's actually going to perform the way you think that they are going to perform before you adopt it and try to put it into your into your environments. I would just add, you know, I think certainly red teaming and You know, scenario-based testing around your models we'll find issues that like you mentioned last week where someone couldn't do instant response because it thought the cyber defense team was actually a hacking unit. But at the same time I think I think back to the last 15 years of cybersecurity, the biggest challenge that we've had is knowing your environment. And so I think step one is really understanding which models are deployed across your environment and you know Shadow AI is the new shadow shadow IT and so really getting control of your boundary and understanding what's in your environment so you really can assess that risk level. Um, all right, let's, let's talk about, so, um. There's no one size fits all model. Um, my, my new favorite term is, uh, tokenomics, and, uh, and so let's talk a little bit about, you know, some recommendations, you know, moving forward, right? You've got the AI, you've got the Chinese models, you've got the U.S. frontier models. Um, how do you, how do you, let's talk model selection, right? How, how do you really, you know, right size, right pick, uh, the, you know, the, the correct model for, for your environment. It's, it's a great, it's a great question, Brad, and, and I, I like tokenomics as well, um, it's sort of the like the biggest headwind of adoption today in, in many ways, um, so, so the, the all models come with a model card, and they have, um, depending on the company or their where their source, there's a lot of valuable information that wouldn't, wouldn't surprise you, you know, traditional benchmarks on accuracy. Different categories for coding, for reasoning, uh, vision, and so forth, um, based on the cybersecurity evaluations we're doing, we're adding diagnostics to standard model cards. There's a couple of areas that we think are very important. One is on the steering sensitivity, which we talked a little bit about today. It's sort of the diagnostic we use if we're going to implement a model into a mission use case, how it perform long term in that use case is different than just hitting an API for a chat, right, much different level of scrutiny and rigor. And the second family of diagnostics is around token efficiency. This is really a measure of how efficient is my model at converting a token to a unit of intelligence. The unit of intelligence will be defined by the particular use case. It could be a back office business operation. It could be an autonomy use case or a cyber use case. So you have to define that as an organization, but it also measures the speed and the cost. And in the cyber landscape, you know, cost and speed matter a lot. You may have a better solution using a low cost, smaller, faster model than a much larger, more expensive model depending on the operation. So we take those diagnostics and others and combine them into model cards, and those model cards sit across our organization where we govern them and use those to select models and route models where we can for different applications. Yeah, token costs are real. Um, it, it, it becomes an issue obviously when you try to operationalize at scale, uh, and so that could steer you towards an openwave model or a closed model based on, on the activity that's performed. I think ultimately what we're saying is also not that open weight models are bad and of course we huge fan of, of, of the Nemotron work that NVIDIA is doing, uh, integrating that. You just have to make sure that once you, you test and evaluate the models that you that. You put into your operations they're going to do and deliver what you think they're going to do at scale again just because it works once doesn't mean it's going to work on time 10. You want to make sure and test rigorously. And once you do that, then it's a matter of which model for which sort of activity elevating to a more frontier model that might be more expensive. But there's also new models coming out every time that you know for an American company, American developer, you know, Meta just switched to closed model and Muse came out. That's very good. Grok 4.5 for cybersecurity purposes and coding also very good. And so you can you can balance these things once you've tested them out. Really get your token costs under control as the context of your company and the context layer adds taxes onto the amount of tokens you'll use. Yeah, I think that, you know, I think that the future will be defined by multi-model environments and you know I think that this whole there's a growing field around sort of model routing based upon the problem. Like if you want to do basic productivity, you probably don't want to accrue the token cost of something like a mythos, right? Um, yeah, I mean, 100% bread, and I think it's, um, how, how we think about that too is, uh. We talked a little bit about the context before and and the prompts, um, also having great uh prompt and context management in an organization it goes hand in hand with model selection and routing. So a large organization like Booz Allen, you can imagine a lot of our employees are asking our enterprise model similar questions. So why route that prompt 1000 times through the same tokens to have a similar response back out? We can cache prompts, do smart things with cash contacts, do smart things with that context to balance out performance and offer additional efficiencies. Let's shift a little bit. You know, one of the buzz buzzword du jour right now is this whole notion of harnesses, and you know, talking to some of the frontier model companies, you know there's different views. Some would say that they're rolling out some of their new models with harnesses, and ultimately their goal is to make them smaller. Others, you know, I've seen Capital One, Visa, open source, cybersecurity harnesses. You know, can you talk? Let's have a conversation around harnesses, you know, what, what directionally, where are the frontier models going, and, and, you know, what role do they really play future from a scale perspective? Yeah, yeah, it's, it's an exciting topic area. So just by level setting kind of the vernacular, you know, a harness to us is the both the framework and the software that serves as the, the interface that connects the large language model to an operational environment. Um, mission system. Um, if the model is the engine, the harness is sort of like the, the, the driver assistance. Um, so it's software and there's a lot of functionality that can be built in that software that turns an API that you would access large language model through, through a real deployable capability, um, things like model routing as, as, as Brad mentioned, um, evaluation, um, capabilities that are monitoring agent and model performance, uh, telemetry collecting data back on inference speed and, and token usage, these are all things that are required to really operationalize. A large language model, um, we, we're finding them to be somewhat domain specific. So for certain use cases there are different skills and tools and capabilities that that we engineer into harnesses. Uh, the very harness that Justin and team created to do our evaluations have unique set of capabilities that are unique to model eval. So, so the harness really is something. That we see a very important spot of the stack now to operationalize this technology. We're even seeing in the cyber front examples where lower cost smaller models with the right harness are outperforming larger models with weaker harnesses. They really do unlock the full potential of a large language model space. And I'll just add harnesses, if you haven't heard about them yet, you're going to hear about them now. Uh, I was telling Brad it's kind of funny, you know, 6 months, a year ago it was all about agentic frameworks and agentic loops and all that stuff, and now it's just totally switched, and the real switch there is that the models themselves become more and more advanced, so it's a matter of harnesses, and these things also include. Not just tool access that they can use but also the guardrails that are there that are present when you want to run a model we've we've done some testing for some of the frontier models and without proper guardrails they hey go look at this file and do an analysis of it and all of a sudden they're looking at the entire directory and telling you what's wrong about it. It's like we never asked you for that. And so that's where harnesses become very important for when you want to operationalize to make sure that you've tested those guardrails. Those are working, the tool calling is up to speed, and then of course token optimization. They all go hand in hand with a robust policy that has, I think, you know. Using the testing that you do to build those harnesses and this is this kind of never ending loop as you plug in new models. So um you mentioned guard rails and testing. We saw some events last week of um some testing that that went awry and caused a pretty significant cybersecurity incident. Talk to me about some recommendations as you're evaluating and testing these models and, and how can you, how can you really buy down the risk there. Yep, that is, that is where experience counts. Uh, I think we've done cybersecurity work at the highest levels as, as a company for a quarter of a century since probably cyber before it was called cyber, and I think the thing that we see from what happened last week was the difference between. Frontier model building and then operationalization, so leaving a model with lower guard rails because you have the ability to do that. Leaving it to do its own thing once you kind of give it an objective and a prompt, not knowing what those were, of course, but the way models will do their own reward hacking is you should have known that. And so it's all about what you can do to limit your risk when you try to do these things, particularly for offensive cyber testing or red teaming of your network. You have to give it very specific timeouts, very specific monitoring of it, the ability to do a kill switch, the ability for it to. You know, not break out, just given the objectives, setting the objectives of the prompts in a way that there's nothing on the other side that it's going to really care about and just really being thoughtful in your process, particularly if you're going to red team your network with these with these models and setting those guardrails in place. All right, let's, uh, let's talk for a minute, um, uh, around open weight models. Justin, you talked a little bit about that and, and sovereign AI. Eric, why don't you start us off there? Yeah, sure, it's, it's exciting to see open weight models, uh, U.S. open-weight models, um, come into the spotlight a bit. There's been a lot of, a lot of great announcements recently, um, from Nvidia. Google and others that are bringing open-weight models to domestic use cases, they provide some advantages in the sense that they can be distilled and they can be fine tuned for specific applications. So rather than have to use an API from a multi-trillion parameter large language model that's closed, we can take an open-weight model and distill it on very unique mission data that makes it highly performing for that particular mission. And make it smaller so therefore it's faster to run, it's cheaper to operate. We can customize it and put it into on an edge device and a sovereign stack on-prem at a data center or an on-prem server. So there's a tremendous advantage in having these available. It's not an either or, right? We want both closed models that have incredible functionality for those frontier use cases. We want the option to use open weight models for things that are a bit more trivial and well understood. Yeah, how about cost? Any, any metrics or measurement around cost savings? So the cost, the cost that you save is in the cost not required to train a full model, and so those costs are not passed to you as the end user. You can pull the model and its weights off of a repo like hugging face or other open repositories. Still pay to run the model, but if you're looking at sovereign workloads where you have already paid for the upfront compute, that's a lot more cost efficient and economical. Some of the standard comparisons are in the neighborhood of 10% operating cost. You also have consistency. So if you if you're building out an AI factory or just an AI capability in your organization, you have a one-time bill for the compute cost and an overweight model which is different than the dynamic pricing of tokens and consumption and so forth. So there's a lot of economic stability. And goodness in that approach. All right, so, um, we're, we're just about ready to shift into Q&A, but let's just kind of wrap up some recommendations. So, so one, I think. You know, we all see it. The world's moving fast. AI models are coming online at an accelerated pace. New releases of both the Chinese and the US frontier models almost on a weekly basis, and everyone ups the game and increases capability. So a couple of recommendations we talked about. One is know your environment, right? Know what models are in your environment, get a grip on shadow AI. We talked about model selection and the criteria to really pick the right model for the right mission. Um, we talked about testing and guardrails in terms of being able to kind of bound and control that within your environment, and we talked about some about the open weight and and AI sovereignty. So good discussion there. Let's let's transition into the Q&A, and I think we have some, we've given the audience, we'd love to hear directly from you in terms of questions. All right, here. Initial question. New one. All right, all right, let, let, let's talk about, here's one, here's a good one. can you unpack adversary steering, uh, with a real scenario and the telemetry you'd monitor to detect it? Yeah, sure. Uh, so we, we talked about prompt steering, and again prompt steering is, um, the behavior of a large language model based on the context it, it observes, not, um, not its underlying underlying weights in adversarial sense, um, most of the time it's the admittance of information or the injection of intentional information. Uh, so think about, uh, a command and control use case or an ISR use case where you're trying to find and collect information. If you're using a large language model, it may intentionally remove and create a blind spot for you in that, in that setting. Some of the things we look for in terms of telemetry are things like inference. Speeds and added latency in our response, things like an unusual amount of tokens being processed based on a request. The response itself can be can be very telling. There's all sorts of information that are under the hood and the reasoning traces that can also be observed that would indicate something is going on underneath the hood of the model that is unusual and stands out. All right, here, here's a, here's a, here's a, here's another question. Uh, did US models exhibit any refusal or bias patterns, uh, that we should watch even if, if less secure? Really love that question. Uh, which, which, which one of you wants to start? I'll start. Yeah, that's a great question. Uh, we did see some refusals, uh, particularly with Claude. I think it was at 2%, so significantly less than than anything else. Most of those patterns were around the general training. Around being able to support, let's say an offensive cyber operation or some sort of military campaign that it didn't necessarily agree with in its in its training writ large, less specific on anything to do with with the US, with the Chinese prompt or with the Russian kind of falsifier that we did. And so we saw that the biases for the US models were actually kind of interesting. They wrote a lot more code because they're most, I'm assuming they were, it was worried more about the actual NIST frameworks and things that it had to go into when you start triggering on a US federal agency as the actual endpoint customer. And so as it wrote more code, it actually did better in writing. The code than the Chinese models which wrote less code so it's kind of like a a bias towards security, which I guess is a good thing in this case, um, but it was just a natural response from that model, which is interesting. Now wouldn't it, I mean, we've seen it. I mean the US models are trained on US philosophy, so I mean there's inherent bias there as well, right? Absolutely. Yeah, I think the difference between the kind of the US philosophy and what you see, there's a lot more in the US the private companies, an anthropic or an open AI will be able to put their guardrails in for things that they deem would be something that could harm a network or that would be inappropriate for their model to do, whereas what you see on the Chinese models is just verbatim quoting of Chinese law, Chinese policy. Didn't really see anybody talk about any sort of like. You know, oh, the Sarbanes-Oxley Act, you know, it doesn't, it doesn't give you any of these sorts of quotes when it comes to the US model, but in the reasoning of the Chinese models, it was always around very specific policy violations, things it wouldn't do based on its uh kind of training after the fact. What I was gonna add, Brad, 111 factor I don't want to be mentioned in the previous part of the discussion that we did control for was language. And so, uh, you know, tokens are roughly 3/4 of a of a word, um, languages like Mandarin and English, they, they influence token usage, um, and inference differently. So that's a factor that was also brought into our test harness to control. All right, next question, uh, from what should leaders tell boards and agency CIOs about the policy posture here? What's prudent today? Anyone want to try? I think first for the CIOs it's it's, it's taking governance. So you know any good policy has to has to stand on a governance capability that is adopted by the organization. We have, we have tried a number of commercial platforms here at Booza. We ended up building our own based on the features and use cases and sort of the preferences of our of our engineers. We have, uh, you know, hundreds of different AI use cases that get that entered and observed by the CIO at the start of every project. And I think um you know based on this research and for many other reasons we we do have a policy where we're using domestic domestic models uh exclusively in our in our workloads across the across the company. Yeah, my, my perspective on this question when you talk to boards and and agency CIOs is I, I think one it starts with education, understanding the risks, understanding the trade space and starting the conversation, um, and really understanding the mission and or the business objectives you're trying to, to, to achieve and mapping the right model to that. I think like we talked about, you know, technology tends to outpace policy. I do fear that we are heading in for a Huawei type moment where you know we watched the technology propagate globally. We made a strategic call to ban that. And we've spent billions of dollars trying to unwind that technology out of our infrastructure. If you look at where we're moving right now with the adoption of Chinese models, you know, hyper accelerating, we talked about 80% of Silicon Valley, a statistic that's been posted, but embracing Chinese models, if we make the call to remove them, it's it's going to be costly. Yeah, I agree. I, I think the one thing I would, I would keep harping on, harp on it again when it comes to the CIO or the C-suite level is within all the different vari..., variables that exist within AI adoption, the thing that you have the most firm in your control and you should not ever let go is making sure that your models are red teamed to align to the activities you want to do, your worldview and your regulatory frameworks. If you give that up because you, you, you don't wanna do the scenario-based red teaming. It's going to end up really bad for you, and I think that to Brad's point, the Huawei, on a separate note, the Huawei model is exactly where we're at now. We've seen this repeated pattern distilling models, creating a photocopy of a cloud frontier model, offering it up for pennies on the dollar, trying to hook companies into that token optimization which is just blatantly taking that. Um, those sorts of things I think are, are where we are, uh, in kind of this geopolitics if you will, between the US and Chinese, uh, model providers. I think, I think we're, we're, we're gonna be quickly entering a time, uh, around, um, return on investment and value and right now, you know, the whole world is out trying to AI enable, you know, different business processes, different mission processes, and I think the, the strategic question will be, OK. You know what model are we using, you know what what what are the tokenomics around it and you know what's the value that that we're getting out of that? And so I think the earlier that you can start that conversation, um you can put a strategy in place to to measure it and move out smartly. All right, next question. What procurement guardrails should program managers add now to reduce model origin risk without slowing delivery? Great question. We, we, we touched on a few, a few relevant kind of parts of this, I think, um. Starting with education and and understanding the use case, um, understanding the choices you have on model selection might be platform selection if you're using a platform for generative software engineering or using a platform for generative data science, they often come with models that they prefer. Certainly on the guardrail space of this use cases dictate the requirements on security in the scenario-based red teaming that Justin mentioned. It also should dictate model selection and token buying patterns. So it would be unusual to subscribe in procurement to a very expensive model for a long period of time without a very defined use case in mind. So a bit more a la carte. And model procurement may make a lot of sense to get off the ground. And once you have enterprise adoption, you feel good about the tokenomics you're receiving, making a longer term commitment after that. When I, when I see questions like procurement guardrails, I think these are really good questions, are a really good question, but I, I immediately kind of shift my mind to. OK, what's the risk and what can be mitigated? What's unmitigatable? We talked about Chinese models producing less secure code while more costly, you know, there are mechanisms to be able to buy down that risk in your DevsOps pipeline, you know, through enhanced scanning and remediation. The risk that I see that that is almost unmitigatable is the inference risk. And what I what I worry about is it's very similar to the open source problem right where you know you're going to be buying large software systems and you're going to be running in system of system environments. The future is going to be like a headless architecture with lots of different agents coordinating and collaborating across an enterprise around different different functions. And the question will be if you know a 2nd, 3rd, 4th tier agent is running a model that's a Chinese model that's doing some level of inferencing, can you actually trust that? And so I think one key recommendation is the software bill of materials, really understand, you know, at a granular level you know what's in the software that you're buying. Yeah, I'd agree with that. I think ultimately the most interesting thing for me in the last year as software development has really become AI software development writ large for anything that's created is that the software supply chain has completely shifted all the way down. So instead of that individual you're hiring, you know what model are you using and then also what is it introducing? I think the other interesting thing that we see from an attacker standpoint, which is relevant to procurement is that you see certain models will reach for certain packages when they go to create software. And you're starting to see a whole lot of attackers out there who are going after specific packages that are probably informed by their use of frontier Model A versus Open-weight Model B and getting in front of the supply chain by hitting these packages, putting in backdoors that then get pulled in to your environment as well as your code that you're creating, and that's just really changed the game and where you see attackers now we're seeing all the way down to individual packages that models select and then they're prioritizing ones that There's there's so much software that I hadn't even heard of two years ago that now everyone talks about because it happens to be a frontier Model's favorite package that it uses for authentication or for vector databases, and it's really interesting to see the usage of some of these things just spike and then the attacks that are happening against those packages. So it really changes procurement and safety of the software supply chain from. The prompt all the way through, it's, it's kind of wild. Um, another question came in around accountability. Um, so where does accountability actually land? Um, and I, I think where they're, where they're going is, in, you know, in an organization. There's a CISO, there's a CIO, there's lines of business, and decisions are being made in all these different organizations. Things are moving fast. The line of business is trying to operationalize technology and move out fast. The CISO is trying to manage enterprise risk, and then the CIO is managing the infrastructure and the IT operations. And so you know who's in charge? It's, it's a, it's a tough, it's a tough space now, right, because there's so, there's, um, a lot of these technologies live in different places. So you may have an enterprise, uh, you may have enterprise large language models being used to support a line of business that is, uh, in a customer's production environment. You have this bridge called a Devsecops pipeline that is connecting these things. So, um, it's a, it's a very tough question. We've had, uh, great conversations with the administration and, and other, um. Other other companies to kind of explore the space. I think our posture is if we're building a system for a customer, uh, we want to be accountable and own the risk associated and the performance associated with the models that are part of that stack, which is why we take evaluation so seriously as part of as part of our build process. Yeah, that's a really interesting question. I think it's the world's moving fast as we talked about. I bring it back to CIOs didn't exist before the Target breach, right? So Target got breached and then who is responsible was basically the C-suite. It was the CEO and the CIO, and so they, you know, created this CIO position, which is probably the most difficult job in all of cybersecurity, which tends to, if there's a breach, they're the ones that get to blame and get fired. I think you. If they host this maybe or if there is an evolution where there is some sort of, you know, chief AI security officer type role that that that's formed because you will start to see breaches happen not obviously from offensive side AI informed attackers but also from the decisions that were made in using AI in your production environment and these sorts of things will probably cause all kinds of chaos in the next. Few months, uh, and I think that that's, that's an interesting question where, where it's gonna land and where responsibility lies. All right, I think, uh, I think we've made it through the, the questions. We have about 5 minutes remaining, um, and so I think we're gonna wrap up. Uh, what I'd like to do is each one of you, uh, why, why don't you give some parting thoughts and then I'll, I'll wrap this up. Justin, all right, I'll start, yeah, uh, again, thanks for joining us. Uh, this, it's really exciting the work that we're able to do, uh, the, the, the partnering that we do across, uh, the frontier model providers and being able to test these things. I think, uh, in addition to all the stuff that we've talked about for US government. Uh, and or defense industrial base customers, I think you know from what we've seen from the deliberate quoting of, of, of CCP policy when you're doing a coding task, I would just say outright if you're in that chain or if I, and not to speak for the government, but if, if I were, you know, relying on a defense tech startup and they were using Chinese models, I don't think I want Chinese policy running on US code for mission purposes. I think that you just got to be able to where you are in industry, be responsible with your model selection based on that knowledge, and there are other cheaper frontier providers as well as really great advancements in U.S. open-weight models and I think it's only going to get better from here on out with the US side on the open weights and just you know adopt those if you're in this sort of line of work where your endpoint customer is the US government. Yeah, I mean, I'll just, I'll just share maybe more of a, maybe a philosophical, uh, closing remarks. So I think, uh, for AI has so much potential and so much excitement and hype around that we're seeing, uh, it transformed businesses like, like cyber is a great example, software engineering, um, chat and research, uh, health. For all of this to work though, um, trustworthiness really is the number one factor. Humans have to trust AI to adopt it and to gain the efficiencies that it's capable of. So this whole topic of evaluating models, really understanding how they work, both closed models, open models, what they can do, how you can tweak them for performance and safety is just such a critical topic. So we are glad to be leading in this space and glad to be sharing these initial results with this group. All right, and thank you both today for your for your insights and your your perspective. Um, we wrote this report to start a conversation. Um, I don't think we have all the answers. Um, we talked about the world changing at a rapid pace um since we actually did the analysis, there's been many, many releases, um, of each of the models we tested and, you know, major new entrants into the market and so this is not a one and done conversation. It's something that we're gonna have to create a consistent dialogue about. I think the reality is that Eric I think you said it best, right? um, you know the future is gonna be based upon trust and you know we need to collaborate together to figure out the right framework, the right conversation, uh, to ensure that. So thank you all for joining us today. Um, we, there's a link, um, on your screen to download the report and, uh, we would love the opportunity to continue a dialogue. Um, you can hit us on LinkedIn. Um, or, um, you know, there's a contact, uh, uh, button on the page. So thank you all and uh look forward to the next conversation. Thank you. 

Webinar: What’s In America’s Code?

Watch Booz Allen’s What’s In America’s Code webinar where our experts unpack the report’s findings and share how to spot hidden risks, strengthen quality gates, and govern AI without slowing delivery.