At the PDF Solutions CONNECT event to be held on October 15–16 in San Francisco, CA. PDF Solutions will present and demonstrate Exensio® Aurora, a new highly scalable architecture for its Exensio analytics solution. Exensio Aurora is purpose-built to handle semiconductor manufacturing data at petabyte scale and to securely deploy agentic AI across manufacturing operations and the supply chain.
In this blog post Said Akar, PDF Solutions General Manager and Head of Development, provides his perspective on four important questions that led to the development of Exensio Aurora.
- How PDF Solutions continues to incorporate in Exensio the latest development in IT technologies?
- Why must we go beyond conventional BI tools in semiconductor manufacturing?
- The value of supporting both online and offline manufacturing use cases?
- How to leverage and enhance what LLM solutions provide in a mission critical industrial application?
Question One
PDF Solutions has 30 years’ experience in semiconductor manufacturing, what is behind the evolution of new architecture of the Exensio Solution leading to Exensio Aurora?
It’s good that we actually use the word evolution because this is what we actually do every few years. Basically, every few years, there is a new combination of problems that need solving at that moment in time, but also a mix of new technologies or opportunities that come along to solve these problems.
This is not the first time we’ve done that. We’ve done that several times in the past. The last one in particular was about data storage and data ingestion. There was a need to be able to handle much, much bigger data sets as far as storage and ingestion. And that was the last level major evolution.
So, what are the needs and opportunities that came along more recently? These really have to do with not just the volume of data that needs to be analyzed, but also, they need to incorporate more automation, more machine learning, and more AI in general in that analysis. And of course, the technology improvements that have been happening recently in these areas have enabled that combination to work.
Meaning now we need to handle, enable our system to handle a lot, lot more data, not just in the storage layers, but also in the analysis, especially semiconductor data because it has very unique shapes and sizes and so on, and a variety of data. But also, to combine that with the ability to automatically do that analysis as well as to incorporate machine learning and parallel processing and AI into that analysis, that combination of need and technology. We felt at the right time how to make this step, this next evolution.
Question Two
Why are conventional BI Solutions not enough to analyze semiconductor manufacturing data?
This is actually a very, very common question because the technologies that are out there when it comes to storing data or analyzing data, it’s not just about the BI tools, by the way, it’s the BI tools and or the data storage type, of technology out there, databases and so on. There’s been a lot of improvements in these over the years, but they always make assumptions about the shapes of the data and the amount of variety of data that might exist within a domain.
What is really special about semiconductor is first, we don’t just have volumes. The volumes are super large, meaning a lot of times our customers tell us that our databases and our volume is the biggest, meaning the semiconductor data is the biggest they have in their organization.
But on top of that, semiconductor is unique because there’s a lot of varieties. So, when you say semiconductor data, it does not mean one type of data. It means many, many, many types that are very different from each other. So, it’s not like one solution works for all of them. You have had multiple solutions within your approach for that, for things to work.
And also, the shapes of these things are different. A very unique thing as an example of semiconductor data shapes is that a lot of times the data is extremely wide but might be very shallow and that gives its own unique, you know, challenges are what they need to be to be handled. So, a generic solution could work.
The bottom line is that you can do a lot, lot better if you are aware of the variety and you understand that variety and you’re aware that the shapes are very different between these different varieties within the same.
Question Three
In the new Exensio architecture how to combine the management of data on the cloud and on-line real time solutions?
This is another very, very important topic for us because we are constantly engaged with customers are trying to implement, you know things like data lake, lake houses, things of that nature that are very highly focused on the analytics on the offline solution part of what is needed. They do not concern themselves too much usually with what’s happening at the edges, the edges far as means at the fab desk, floor, assembly floor, things like that. And that’s from our perspective, always a big weakness of such situations.
The reason of course is that by having an integrated system that can manage not just the offline analytics and your data lakes or data warehouses but also manage the edge solutions in an integrated fashion has huge advantages. Why? What’s happening at the edge usually is a combination of data collection and control. And both of these can be improved on themselves by integrating them with the analytics to offline analytics and vice versa. The offline analytics can be also improved by being integrated with these two.
As an example, data collection. What’s the most important thing there from the analysis point of view is data quality. So, is the data complete? Can it be more complete? Is it correct? Is it validated? All these things can be better if there’s presence of the same system in an integrated fashion at the floor. That’s for the fab because then you’re able to prove that data collection basically it improves the analytics, can improve the data collection and the control, the data collection and control in an integrated fashion, the analytics can improve the analytics.
Examples of that would be let’s say the control. Control can benefit dramatically from a feedback loop capability, meaning what I learned from the analytics, I’m able to deploy better control. So, my models that I build from the office analytics, I can improve on those models and then I can prove what’s happening control wise at let’s say or desk floor and same thing. The analytics is highly dependent on data quality. By being at the data collection step, it means I can collect more data, more variety of data, better data, meaning in a sense it’s better-quality data, more complete data and so on.
An integrated approach between data collection control at the edge and the off analytics is actually a crucial part of our system and it’s usually we look at it as a mistake not to have such an integrated view. All of these things are possible without an integrated system, but much, much harder to accomplish.
Question Four
Machine Learning has been used for a long time, in Exensio how do we support a range of AI approaches including LLMs?
Of course, LLMS are getting so good. Everybody’s tempted to say, OK, an LLM can do everything for me. And of course, everybody agrees LLM are already very good.
They continue to improve. There are typically exponential improvements, and they are expected to improve a lot more. So, the main question is not that how capable they are today, or how capable will LLMs be in a month from now or six months from now. But the question is always can you improve on it as a vendor, in particular, can you improve on an LLM irrespective of what the status of it is at that moment or how much it will, how good it will be a month from now.
And our answer to that question is yes, there are certain things that are inherent to the way LLMS work, and that based on those things that are inherent to it, there are ways to improve it, no matter how good it has gotten at a certain point. What is meant by that are a couple of things. LMMs will hallucinate. But yes, it can be improved on. It can be reduced over time because they’re getting better, but they are not deterministic systems. It means in the end there is potential for hallucination. So, you want a system that uses LLMs but that is capable of controlling that. Hallucination, in a sense, is a type of creativity. You want creativity in your system, but not everywhere.
Sometimes you want to be much more deterministic in your approach. One area we can always contribute to is to determine where creativity is needed versus not and how much of it is needed versus not. And that’s something that knowledge of the domain and having things in the system that enable that kind of control is important. So that control on the creativity part is important. And this usually requires not just again, a generic solution, it requires domain knowledge to where such creativity needs to be controlled to be more of it or less of it, and how to do it from the main and knowledge and expertise.
In addition, LLMS are notorious for basically not necessarily easily enabling repeatability and visibility and things like that. Meaning if I ask an LLM certain task, it might do it one way today, tomorrow it might do it another way. Or maybe if I use another version of the LLM or another flavor of the LLM, the answer might be quite different. These are things that sometimes are OK, but a lot of times when you ask the same question the next day, you should get the same answer ideally if it’s the correct answer.
So, there’s things like that. Also, it does not need to be a black box. We need to be able to say this is the result that was generated. This is why it was generated and why it is correct and possibly even easily edit such results. So, repeatability, editability, visibility of what’s being created is super important and we believe we are able to do that.
And the last point usually is that there is a lot of domain knowledge within any domain, especially in semiconductor that is not necessarily out there available to a generic LLM.
That context can be made available. Several flavors of customizations of LLMS can be made available to these LLMS. So that combination of customizing LLMS and the ability to make them repeatable and more visible and so on, which effectively then adds memory to the system is, is, we think necessary independent of how advanced that any particular LLM gets.
More information and registration for PDF Solutions CONNECT 2026 can be found at this link PDF Solutions CONNECT
PDF Solutions CONNECT provides a unique opportunity to meet with the team that developed Exensio Aurora and to hear from customers who were early adopters of some of its capability.
Visit PDF Solutions CONNECT 2026 conference website to learn more about the agenda, speakers, location, logistics and registration.