FabOS: A Fabrication Operating System
FabOS applies AI and machine learning to semiconductor manufacturing: the wafer fab and the advanced materials and critical components that supply it, from specialty chemicals and CMP slurries to high-purity filtration. It catches process drift while values are still in spec, flags at-risk batches when a fault occurs, predicts yield and metrology, and answers engineers’ questions in plain language.
FabOS is our reference architecture, not an off-the-shelf product. It packages patterns we have proven in production, including large-scale traceability for a semiconductor materials manufacturer and Graph Data Science for at-risk batch analysis at Entegris, and we build it inside your environment around your tools and data.
Sources
- Tool logs & sensor traces
- MES lot & unit lineage
- Metrology & SPC
- Material & supplier genealogy
Capture & bind
- Immutable raw files
- One parser per format
- Confidence-scored binding to runs
FabOS core
- Knowledge graph: lineage & context
- Columnar store: traces & maps
- Governed semantic layer
Outcomes
- Drift & anomaly detection
- At-risk batch & fault impact
- Virtual metrology & yield
- Natural-language answers
1. The Problem: Fab Evidence Lives in Silos
Every unit of manufacture leaves a trail, but it is scattered. Per-run sensor traces sit on tool PCs in each vendor's own format, often owned by no one. Lot and unit lineage lives in the MES or batch records. Spec limits, SPC and detailed metrology live in an engineering database. Upstream, the materials that went into each batch are tracked in ERP genealogy records. Answering a simple question such as "why did this unit perform so well?" means stitching these together by hand.
FabOS joins them once, at the right level of identity, so an engineer sees lineage, recipe setpoints, what the tool actually delivered and how the unit measured, side by side.
2. What FabOS Delivers
- Drift and anomaly detection: statistical and machine-learning models flag unusual runs and slowly drifting tools while values are still in spec, before they show up as scrap.
- At-risk batch and fault impact analysis: when a tool, step or material lot is found to be faulty, FabOS identifies every downstream batch it touched so containment is fast and scoped.
- Virtual metrology and yield prediction: deep-learning models on sensor time series and inspection maps estimate how a unit will measure and perform, with less physical sampling.
- Faster root cause: each result is compared against matched control runs, so engineers spend less time assembling evidence and more time fixing the cause.
- Plain-language answers: AI agents let engineers ask questions about runs, tools and lots without writing queries.
3. How It Works: A Knowledge Graph Foundation
Every one of those outcomes depends on context: which unit ran on which tool, with which recipe and materials, and how it measured. FabOS keeps that context in a Neo4j knowledge graph organized in four layers. The graph holds identity, lineage, context and precomputed statistics; high-volume arrays such as sensor traces and inspection maps live in a columnar Parquet store, with every file keyed back to a graph node.
Product-Agnostic by Design
The model is built around a generic unit of manufacture, not a wafer. The same lineage, measurement and decision layers describe very different products; only the unit type and its properties change:
Built to Production Standards from Day One
- The tool is never touched. Capture runs beside the tool's control software, never inside it.
- Raw data is immutable and content-addressed, so every analysis can be reproduced from source.
- Read-only, off-peak access to production MES and engineering databases.
- One canonical parameter per physical quantity, however many vendor formats report it.
- Every link carries its evidence. Each trace is bound to its process run with a method and a confidence score, and low-confidence links are never used to set norms.
4. Proven in Production: Traceability and Fault Impact
FabOS builds on two production deployments that trace lineage in both directions: backward from a finished unit to its sources, and forward from a fault to everything it affected.
Supply-Chain Traceability: From Finished Unit to Source Lot
For a semiconductor materials manufacturer, we delivered a knowledge graph powering a Digital Product Passport across multiple manufacturing lines, unifying ERP batch genealogy, bills of material and quality data in one traversable model.
- 80M+ genealogy relationships and 250M+ sampling records, with tens of millions of quality measurements.
- Sub-second traceability forward and backward: serialized unit → batch → sample → measurement → source material lots.
- Built for root-cause analysis, recall containment and regulatory compliance, with measurements classified as key process inputs and outputs (KPIV/KPOV) for process control and yield work.
At-Risk Batches: From Fabrication Fault to Impact
For Entegris, we integrated Graph Data Science into the fabrication process to identify at-risk batches and support impact analysis when fabrication faults occur. When a fault is found, the graph shows which batches were exposed, so teams can contain the problem quickly without holding back unaffected product.
5. Ask in Plain Language, Get a Governed Answer
FabOS provides a natural-language interface built on AI agents. Engineers ask questions the way they would ask a colleague, from the details of one fabrication run to fleet-wide comparisons:
“Show the full path, recipes and measurements for wafer W-047.”
“Was today's run on Reactor A normal? Which channels made it unusual?”
“Is any mass-flow controller drifting or getting noisier while still in spec?”
“Which wafers went through Etch-02 after its last maintenance, and how did their CD compare with the cohort?”
“Which slurry batches used this abrasive lot, and how did their particle-size distribution compare?”
“Show the test history for filter F-1182 and every membrane and resin lot it contains.”
“Trace this finished unit back to every source material lot it used.”
The agent does not improvise database queries. It fills a structured comparison (subject, measure, baseline, statistic and eligibility rules), and the semantic layer generates a bounded query. If the comparison group is too small, the answer is "insufficient cohort", never a number computed on too few runs. Every result is saved as a versioned, replayable finding with citations back to the underlying data.
6. From Questions to Analytics and Deep Learning
Answering questions is only the first use. The same governed layer turns fab data into analysis-ready datasets:
Statistical Analysis
Control limits and Cpk, residuals against recipe setpoints, multivariate Hotelling T², and rolling-cohort drift that catches slow shifts while values are still in spec.
Feature Workbench
An analytics environment for developing features: phase-level dynamics such as overshoot and settling time, timing and lineage features, and graph features such as embeddings and community structure.
Deep-Learning Extracts
Versioned, point-in-time Parquet snapshots with every column keyed back to the graph, ready for virtual metrology, yield prediction, and deep-learning models on sensor time series and inspection maps.
Models follow a governed loop: explore variance across steps, tools and recipes; diagnose root causes against matched control runs; optimize with recommendations that pass a human approval gate and never write to a tool; and operationalize with registered models, drift-triggered retraining and engineer feedback captured as labels.
7. Phased Delivery
FabOS is designed to start small and grow without rework. The first phase uses the production keys and contracts, so nothing is thrown away.
- Logger first — capture and baseline tool runs from one or two tools; every run gets a cross-run verdict within a day.
- Fab lineage — bind tool runs to MES lots and units with confidence-scored links.
- Metrology — join tool features and measurements per unit; report the first virtual-metrology correlations.
- Analytics & natural language — plain-language questions, persisted findings and ML-ready extracts.
Engagements typically begin with a proof-of-efficacy study (about 12 weeks, with success criteria agreed up front) and progress to a durable platform build (about 5–6 months). Everything runs inside your own cloud environment and inherits its security posture.