That is why design teams, verification teams, and implementation teams are coming to us now and asking about token efficiency, token budgets, and how to budget for this.
The next step is that the token budget, wasted or not, is becoming a big enough percentage of those mask costs that it’s starting to matter.
They could break it down—cost per engineer, cost per chip —and that’s how they projected their budget.
This is another way to bring down the token cost.
Those are the ways you can further reduce the token cost.
Key Takeaways:
Discussions have shifted from unlimited token budgets to restricting tokens or dollars at the engineering level, and rising costs are driving interest in open-source models.
Human engineers need tool literacy and workflow literacy to effectively enable agentic AI across the foundation model layer, tools and memory layer, and orchestration layer.
Reinforcement learning, flow capturing, and model routers are all needed to ensure agents do the precise workloads required, especially in a mixture-of-experts system.
Experts At The Table: Semiconductor Engineering sat down to discuss recent developments and challenges when using agentic AI for chip design, with Matt Graham, senior group director of verification software product management at Cadence; Harrison Balistreri, head of business development and strategic partnerships at ChipAgents; Alexander Petr, senior director and portfolio manager at Keysight EDA; Sathish Balasubramanian, head of product for EDA AI at Siemens EDA; and Anand Thiruvengadam, executive director and head of AI product management at Synopsys. This roundtable was held behind closed doors at the recent Design Automation Conference. What follows are excerpts of that discussion. This is the second of a three-part series. Part one is here.
Fig. 1: L-R: Synopsys’ Thiruvengadam; Keysight’s Petr; Cadence’s Graham; ChipAgents’ Balistreri; and Siemens’ Balasubramanian.
SE: When design teams consider the return on investment for using agentic AI tools, or give customers an estimate for design costs, they must now factor in cost per token. What are you seeing?
Graham: In the setup for agentic AI, nothing is free —literally, theoretically, or tangibly. That is why design teams, verification teams, and implementation teams are coming to us now and asking about token efficiency, token budgets, and how to budget for this. ‘How do I budget for deploying your tools, deploying the tools that your tools are going to interact with, and deploying the generic coding assistants?’ Even in the last three to four weeks, that has started to feature strongly in the conversations about agentic AI. Everybody in the industry at large agrees there’s value here. All the folks in this room have demonstrated, one way or another, that we can provide ROI to some degree at varying levels. The question is, how do we genuinely measure that ROI, and how good is any one of us getting at token efficiency and the ability to prove that we can give the right answer, at what cost, and how efficiently?
Petr: Before we go down the path of token ROIs, we need to think about how many iterations it takes to get a chip right. That’s where trust and verification come in. We can talk all day long about designers trusting the process and taping something out. If a TSMC mask set costs you multi-million dollars, or hundreds of million dollars, there is no such thing as trust. There’s only trust and verify, and that’s where the issue comes from. Don’t forget that verification happens with physical simulation engines. That’s where we talk multiphysics. There’s no such thing as just adding an ontology layer and having LLMs reason and verify, while thousands of agents waste tokens to double, triple, or quadruple-check every decision. You run engines at the end, and if you get the thermal budget wrong, it doesn’t matter what your digital design or analog design looked like. Your IC is going to burn, and you’re dead in the water. If you got the power wrong, it’s going to burn. So you’re going to run multiphysics across the whole board, up and down, with the engines with the capabilities of doing real verification. That’s not a human looking at code and saying, ‘Oh, there are 20,000 lines of code. I don’t understand. Let me run my agent on it and let my agent decide if it’s a good or bad thing.’ That’s what we’re seeing in the software engineering world. Then Amazon ships an update, and the AWS server goes down because no one really understands what was coded anymore. That is the new AI slop in the real world. We see that in chips as well, where we now enable agents to do chip designs that humans wouldn’t design. We’re trying to use parasitic effects as part of a feature to get chips out. To play devil’s advocate here, we can streamline all of this with fancy AI, super-agents, agent swarms, and start wasting tokens left and right. But if the chip or the mask head is wrong, if anything goes wrong at that level, you’re going to run another iteration, and that’s going to cost you more than any of your tokens. That is still the threshold we have to live by, and especially if you’re talking about the new big CPUs, the AI chips. They’re huge. You get this wrong the first time, you’re dead in the water. Your company most likely goes belly up.
Graham: The point is that we know how to put out chips. We’ve proven that AI can accelerate that to a certain degree. It doesn’t remove the requirements. The next step is that the token budget, wasted or not, is becoming a big enough percentage of those mask costs that it’s starting to matter. Last year, nobody cared.
Petr: The real problem we’re facing now is that if you look at how design companies work, in the past, they came up with a project proposal, they did the math, and they said, ‘I need to sell this chip for X amount of dollars, and I need to produce a billion or trillion of those to be profitable.’ Then you do a cost breakdown structure all the way down to engineering cost and software cost. EDA budgets used to be fixed. They’re not fixed anymore. That creates the issue of how to estimate the chip price when you have a ballooning budget at the bottom that you can’t control. In the past, the customers negotiated really hard with all of us about pricing, and we settled on a cost. They could break it down—cost per engineer, cost per chip —and that’s how they projected their budget. They decided whether it makes sense to build that chip or not. Now we’re coming in and saying you can be way more efficient. The value proposition is you do 10 chips instead of two. Same engineering, great. But now we have a token budget, which balloons out of proportion.
Balasubramanian: The economics go out of scale.
Petr: What happens is every CFO just opens up their wallet because they don’t know what else to do. There is no such thing as control at this point. My prediction is that next year, that discussion will be completely different. We will see token efficiency as the main issue. We will see customers trying to go with SLMs (small language models), and they will try all kinds of ways to limit token expense because they need costs to be predictable.
Thiruvengadam: It’s already becoming a reality. With many of our customers, the discussions have shifted from unlimited budgets to restricting, at the engineering level, a certain number of tokens or dollars. That’s happening, and that’s also why there’s a trend toward open-source models. People are really excited about ‘frontier’ open-source models.
Petr: If you want CapEx that is predictable, which you can depreciate over a certain number of years, put a frontier model on it. The whole discussion of where it comes from is kind of under the table. But that’s happening right now.
SE: What difference does a particular choice of AI model or LLM make in terms of cost per token when using agentic AI in the chip design flow?
Balasubramanian: Going back to the question of what engineers should control, one thing we demonstrated with Nvidia for long-running agents is the routing for LLMs. That’s on the tool-provider side, or who’s setting up the flow. What we are seeing is that customers are getting much smarter. They are not giving out the latest and greatest models, or any frontier models, just open-source models for the initial phase of design, or designs that don’t really need them. They can communicate and do 98% of what anyone can do. It is just the 2% very extreme [part of the design flow] that needs the greatest models. Six months ago, everyone was using a Ferrari for their daily grocery run. That doesn’t make sense at all. People are figuring out where to use AI products, and which specific LLMs to use. For example, Switchboard (an AI-powered communication platform) and Nvidia have a very good way of routing LLMs. It’s on the designer side, on the provider side, and on the software side to figure out which LLMs work best for a particular workflow. Previously, we only thought about how to get to the answer fast. Now, the cost comes into play. If it is open source, customers love it. They say, ‘Okay, I can air-gap my system. I can be on-prem. I can encrypt. I can do a lot of things.’ That’s where things are going.
Petr: A fun fact here is that some companies do the math and say, ‘Hey, an engineer is cheaper than having an agent running for the same job.’
Balistreri: It depends on your agent. We see tokens per engineer per year going to 100 billion in the next 18 months, based on trends in how people use our products. I’d say it’s not a wasted or inefficient spend. For our customers, this budgeting and ROI conversation is a reality today. It’s not coming next year. It’s right now, and the new baseline. If you look at what we’ve done in our bespoke loops, side by side with the same baseline, we are 40% to 50% cheaper on the same tasks. Using the same underlying Opus model, it’s the same thing. That’s why I think open source is rising in the conversation. But we do a ton of benchmarking. The models don’t perform on these complex tasks without significant fine-tuning and post-training, which we do. We released our own host-trained, open-weight model that performs similarly to frontier models on chip-specific tasks with long horizons and difficult SoC-level tasks, but it’s 5X cheaper. Customers are pushing for this because it needs to be deployed on-prem. It needs to be deployed securely — your design and your flows.
I’ll add one more layer. There’s local context that most agents can use in generic LLM schemes, but what you need is the global context. You need what is unsaid and unspoken between teams and individuals. We do a ton of work going into organizations and mapping this into flows so they can get ROI through efficiency. It’s both at the agent loop and harness level. It’s at the organizational flow level, and then it’s the model itself. Right now, the Kimi model (K3 AI by Chinese startup Moonshot AI) delivered a chip at $48. I heard someone say that’s like making a car that can only go one mile an hour. It’s important to note that some of the engines and tools are just software. They will be subsumed into the models, especially as we work with the foundries to accelerate that. But for now, to deliver full-flow autonomous outcomes with true ROI, you need the ontology, the flow, the harness optimizations, and the underlying [model] that customers can deploy on-prem and not pay these ridiculous token prices for a frontier model.
Petr: Everything you are talking about is basically flow capturing — how to extract the flow from an organization. You are now talking about creating designated skills that are hopefully deterministic, creating tools that encapsulate prior, non-automated pieces of engineering knowledge that only live in engineers’ heads. A ton of work goes into enabling those glue pieces between the big tasks, and that has nothing to do with any frameworks or LLMs. You can take any model you want. The more time you invest in enabling those things, the more capable any LLM becomes. So yes, you can use Kimi, but you may need to spend more time enabling your workflow and making sure you capture all the nitty-gritty details. If you use a big Ferrari, it flattens the waves and rolls over some of the things that are unspoken. But it can be done. It’s just a matter of where you put your effort. It’s still possible to go with Opus and spend a little bit more on tokens. Or, you just invest more time in enablement.
Balistreri: I don’t think it’s just enablement. It’s also fine-tuning the model itself, and there are certain tasks that smaller models just can’t handle. If you think about who your partner should be to implement agentic AI and produce ROI with full-flow autonomous outcomes, you need someone who can go into enterprise and map the flows, but also deliver you a model — if it’s an open-weight model — that is post-trained potentially on your data, on your chip designs, on your libraries, or brought up from the base model performance. Because base-model performance, even for larger open-source models, just isn’t there yet. You want to make an agentic AI bet. Who’s going to help you own your intelligence? Who’s going to help you not give away your organizational alpha? And then who’s going to produce that ROI? Right now, you can’t just take a model out of the box and do it if you’re trying to compare with the frontier models.
Graham: What you’re saying is reinforcement fine-tuning, or something similar to it. I would argue that’s not the only way. There’s still a place for the frontier model, and if you think software can be subsumed into the model, fine-tuning probably can be subsumed into the model more easily. It will be very interesting to see [what wins] in another year. Is it fine-tuned models? Is it a bifurcation of the usage between hard and easy problems? Is it some combination of the three of those? How much air gap is required? These are all questions that have no clear answers. It will be some combination of all of them. In terms of setup before you deploy a model, how you do this is a real challenge for enterprises right now, and some of them are making bets in every one of these areas. Others are taking selective bets. What does that infrastructure need to look like? Is it an AI partner building custom models? Or are they building their own custom internal models, depending on their capability? Is it infrastructure that enables addressing many language models, both on-prem and off-prem? There are so many different choices.
Thiruvengadam: Tying back to the original seed of this discussion, which is token efficiency, everybody’s saying the right things. Context intelligence can help alleviate some of those token inefficiencies, or rather, reduce them. Tool literacy and workflow literacy can also be important. This is another way to bring down the token cost. These are things you can do easily. The engineer is literate in the tools, and that literacy needs to transfer to the AI systems so they become more efficient at calling LLMs, for example. That’s determining the budget. So tool literacy, workflow literacy, the whole context intelligence — that’s one way. In terms of the model discussion, the point is not that, ‘Hey, you can use an open-source model to solve most complex problems.’ The point is that it opens the possibility. If you have a highly capable open-source model, it gives customers and vendors alike optionality. They can consider taking a highly capable open-source model and fine-tune it. Then, for the mixture of experts, it’s just model routing. You can deploy a model router. We can do that. We are doing it already. We can deploy a model router that looks at the nature of the task and its complexity, then decides which model to use. Those are the ways you can further reduce the token cost.
Petr: Maybe we can summarize it very simply. Some companies mix agentic frameworks into the LLM, so it becomes harder to distinguish what is the LLM and what is the reasoning piece. But the LLM itself is the least important thing in the whole stack nowadays. Everything else, including the enablement — and there are so many technologies that sit on top, which give you token efficiency, context, all the references, the tool chains to scale — that is where the real work lies. Open models are six months behind frontier models. Six months is infinity right now. If you ask me today, ‘Do I want the open-source model from [six months in the future], I would say, ‘Hell yeah,’ because I know it’s going to be a decade further in the future. At some point, the stuff will saturate, and everything will be in the enablement we do.
Read part one of the discussion:
Agentic AI Success Relies On Excellent Human Scaffolding
Before deploying agentic AI into a chip design, engineers must define the domain ontology and agentic harness. Fundamental engines are still crucial.