Chapter 1: Why Process Informatics?
This chapter articulates the bottlenecks in process development (multiple variables, scale-up). It captures the value of data-driven optimization through a big-picture view.
💡 Supplement: In the field, designing "what and how much to measure" is essential. Starting from data visualization brings the challenges into sharp relief.
Learning Objectives
By reading this chapter, you will acquire the following: - Understand the historical evolution of chemical process development (from ancient times to the present) - Be able to explain the limitations and challenges of conventional process development - Understand the social and technical background that makes PI necessary - Learn from concrete examples of chemical plant optimization
1.1 The History of Chemical Process Development: Thousands of Years of Trial and Error
The development of the chemical industry is closely tied to the evolution of human civilization. Many of the products we use—pharmaceuticals, plastics, fuels, textiles—are produced through chemical processes.
Ancient Times: The Age of Rules of Thumb
Distillation Technology (around 3000 BCE)
Distillation, one of humanity's oldest chemical processes, was used in ancient Mesopotamia for perfume production. However, this technology was based entirely on empirical rules of thumb.
- Development method: Trial and error, and oral transmission across generations
- Development period: Optimal temperatures and times discovered over hundreds of years
- Knowledge accumulation: No records, only the tacit knowledge of artisans
Fermentation Technology (around 7000 BCE)
The production of beer and wine is the oldest bioprocess, harnessing the metabolism of microorganisms. Although ancient people were unaware of the existence of microorganisms, they empirically discovered the appropriate conditions (temperature, sugar concentration).
The Modern Era: The Beginning of the Scientific Approach (1800s-1950s)
The Haber-Bosch Process (1909)
German chemists Fritz Haber and Carl Bosch developed an industrial process to synthesize ammonia from nitrogen and hydrogen. This was the beginning of design based on scientific principles.
- Development method: Theoretical calculations based on thermodynamics and reaction kinetics + experimental verification
- Development period: About 10 years (from laboratory scale to industrial plant)
- Impact: Mass production of chemical fertilizers sustains the world's population
- Nobel Prize: 1918 (Haber, Chemistry), 1931 (Bosch, Chemistry)
However, discovering the optimal catalyst (iron-based) still required an enormous amount of trial and error. Haber is said to have tested about 6,500 candidate catalysts.
The Development of the Petrochemical Industry (1940s-1970s)
After World War II, the chemical industry using petroleum as a feedstock developed rapidly. From basic chemicals such as ethylene and propylene, plastics such as polyethylene and polypropylene came to be mass-produced.
- Development method: Chemical engineering theory (mass balance, energy balance) + pilot plants
- Development period: 5-10 years to develop a new process
- Scale-up challenge: Differences in conditions between the laboratory and the plant (reactor size 1,000 times larger or more)
The Present Day: The Age of Process Control (1980s-Present)
Introduction of DCS (Distributed Control System) (1980s)
Distributed control systems made it possible to automatically control the temperature, pressure, and flow rate of the entire plant. This significantly reduced the burden on human operators.
- Effect: Improved safety, stabilized quality
- Challenge: Optimization still depends on human judgment
Smart Factories and Industry 4.0 (2010s-Present)
Through the integration of IoT sensors, big data, and AI/machine learning, real-time optimization is becoming possible.
- Characteristics: Collecting data from thousands of sensors every second
- Goal: Minimize human intervention and autonomously maintain optimal operating conditions
- Challenge: How to leverage the data?
Challenges Revealed by History
Looking back over thousands of years of chemical process development history, the following challenges emerge:
- It takes time: 5-10 years to develop a new process, and several more years for scale-up
- It is expensive: Hundreds of millions of yen to build a pilot plant, tens of billions of yen to build a full-scale plant
- Dependence on trial and error: An enormous number of experiments are needed to discover optimal conditions
- Dependence on tacit knowledge: The experience and intuition of skilled operators are indispensable
A question: What if we could systematically optimize chemical processes with data and AI?
This is the starting point of Process Informatics (PI), the next-generation process development method.
1.2 The Limitations of Conventional Process Development
Modern chemical process development has become far more sophisticated than in ancient times. Nevertheless, it still faces major challenges.
Challenge 1: It Takes Time
A typical process development timeline
Years 1-2: Literature review and laboratory-scale experiments
↓
Years 3-5: Bench-scale experiments (100 mL-1 L scale)
↓
Years 6-8: Pilot plant construction and operation (10-100 L scale)
↓
Years 9-12: Scale-up optimization
↓
Years 13-15: Commercial plant construction and trial operation (1,000-10,000 L scale)
↓
Years 16-20: Start of commercial production and continuous optimization
Result: Commercializing a new process takes 10-15 years on average [1,2].
Challenge 2: It Is Expensive
Typical costs of process development (medium-scale chemical plant)
| Development stage | Cost (yen) | Period |
|---|---|---|
| Laboratory research | 50 million-100 million yen | 2-3 years |
| Bench scale | 100 million-300 million yen | 2-3 years |
| Pilot plant | 1 billion-3 billion yen | 3-5 years |
| Commercial plant | 10 billion-50 billion yen | 3-5 years |
| Total | 11.6 billion-53.4 billion yen | 10-16 years |
Because 50-70% of projects never reach commercialization, the cost of failure is also enormous.
Challenge 3: The Difficulty of Scale-Up
Problems in scale-up
Chemical reactions behave differently depending on the size of the reactor. This is called the "scale-up effect."
Differences between the laboratory (100 mL) and the commercial plant (10,000 L):
| Factor | Laboratory | Commercial plant | Effect |
|---|---|---|---|
| Mixing efficiency | Uniform | Non-uniform (dead zones exist) | Reduced reaction rate |
| Heat removal | Easy | Difficult (small surface-area-to-volume ratio) | Temperature rise, increased side reactions |
| Residence time distribution | Uniform | Non-uniform | Reduced yield |
| Scale ratio | 1 | 100,000 times | Hydrodynamic behavior changes |
A concrete example: a failed scale-up of a pharmaceutical process
- Laboratory scale (100 mL): 95% yield
- Bench scale (1 L): 90% yield
- Pilot plant (100 L): 78% yield
- Commercial plant (10,000 L): 65% yield
Cause: Because it was an exothermic reaction, heat removal could not keep up at large scale, and side reactions increased.
Time required to address it: Two years of additional experiments and 300 million yen of additional investment.
Challenge 4: Batch-to-Batch Variation
The problem of quality variability
In chemical processes, even when you intend to operate under the same conditions, product quality varies from batch to batch (production lot).
Causes of variation: - Variation in raw materials (purity, impurities) - Changes in environmental conditions (outside temperature, humidity) - Aging degradation of equipment - Differences in operator skill
A concrete example: pharmaceutical manufacturing
Pharmaceuticals require strict quality specifications. Under the standards of the FDA (U.S. Food and Drug Administration): - Active ingredient content: within the range of 95-105% - Impurities: <0.1%
Conventional management methods: - Record the measurement results of each batch - Discard batches that fall outside specifications (cost loss) - Weeks spent investigating the cause
Problem: Reactive, after-the-fact responses cannot prevent quality defects.
Challenge 5: Dependence on Experience and Intuition
The tacit knowledge of skilled operators
In chemical plants, the "intuition" of skilled operators is important:
- "When you hear this sound, there's a problem with the reactor's mixing."
- "If it's this color, the quality is good."
- "Since the humidity is high today, lower the temperature by 2°C."
Such tacit knowledge is extremely valuable, but it has the following problems:
- Difficult to systematize: Because it is based on personal experience, it is hard to put into words
- Reproducibility problem: Even under the same conditions, judgments differ from person to person
- Time to train young workers: Becoming skilled requires 10-20 years of experience
- Loss of knowledge through retirement: When veterans retire, the expertise is lost
A question: What if we could convert tacit knowledge into explicit knowledge with data and AI?
1.3 Case Study: A Real Example of Chemical Plant Optimization
Let's take a detailed look at a catalytic reaction process optimization project at a major chemical manufacturer. This is a typical example illustrating the difficulty of conventional methods and the potential of PI.
Background: The Challenge of Improving Yield
Product: A fine chemical product (a pharmaceutical intermediate) Process: A catalytic reaction (three reaction stages proceed within the reactor) Challenge: The yield had stalled at 70%, while the industry leader achieved 85%
The demand from management:
"Improve the yield to 78% or higher and increase annual sales by 500 million yen. The deadline is 12 months."
Phase 1: Optimization by Conventional Methods (6 Months)
Approach: Trial and error by experienced engineers
Variable parameters: - Reaction temperature (80-120°C) - Pressure (1-5 atm) - Catalyst amount (1-5 wt%) - Residence time (30-180 minutes) - Feed flow rate (10-50 L/h)
Experimental design: - Each parameter varied at 3 levels - Total combinations: 3^5 = 243 - In practice, 50 conditions were selected based on experience
Experimental schedule: - 1 week per condition (preparation, reaction, analysis) - 50 conditions × 1 week = 50 weeks (about 12 months)
Result (at the 6-month mark): - Number of experiments completed: 24 conditions - Highest yield: 73% (+3 %pt) - Would not meet the deadline
Phase 2: Introduction of the PI Method (3 Months)
By management's decision, the project was handed over to the PI team.
Step 1: Collect existing data (1 week)
Operating data from the past 5 years was collected: - Process parameters: temperature, pressure, flow rate, catalyst amount (every second) - Product quality data: yield, selectivity, impurities (per batch) - Total amount of data: about 500 batches, 1,000,000 data points each
Step 2: Build a machine learning model (2 weeks)
Model used: Random Forest
# Simplified code example
from sklearn.ensemble import RandomForestRegressor
# Descriptors (inputs): temperature, pressure, catalyst amount, residence time, flow rate
X = process_data[['temperature', 'pressure', 'catalyst', 'residence_time', 'flow_rate']]
# Target (output): yield
y = process_data['yield']
# Model training
model = RandomForestRegressor(n_estimators=100)
model.fit(X, y)
# Prediction accuracy
print(f"R² score: {model.score(X_test, y_test):.3f}")
# Output: R² score: 0.893 (high accuracy)
Model performance: - R² = 0.893 (89.3% prediction accuracy) - MAE (mean absolute error) = 2.1% (yield prediction error of ±2.1%)
Step 3: Exploration by Bayesian optimization (2 months)
Using the machine learning model, optimal conditions were efficiently explored.
Number of conditions explored: 20 conditions (the conventional plan called for 50 conditions)
How Bayesian optimization works: 1. The model proposes the "next condition to try" 2. Verify it with an experiment 3. Add the result to the data and update the model 4. Repeat
Result (at the 3-month mark): - Highest yield: 85.3% (+15.3 %pt, exceeding the 78% target) - Optimal conditions: - Temperature: 102°C (previously 95°C) - Pressure: 3.2 atm (previously 4.0 atm) - Catalyst amount: 2.8 wt% (previously 3.5 wt%) - Residence time: 65 minutes (previously 90 minutes)
Phase 3: Verification at the Actual Plant (1 Month)
Pilot test: Confirmed the predicted yield of 85%
Rollout to the commercial plant: - First batch: 84.2% yield - After stable operation (10 batches): average yield 85.1%
Additional findings: - Energy consumption reduced by 30% (the effect of lower temperature and shorter time) - Byproducts reduced by 40% (improved selectivity)
Results and Impact
| Metric | Conventional | After PI | Improvement rate |
|---|---|---|---|
| Yield | 70% | 85.3% | +15.3 %pt |
| Development period | 12 months (planned) | 4 months (actual) | 67% shorter |
| Number of experiments | 50 (planned) | 20 (actual) | 60% fewer |
| Energy consumption | 100 (baseline) | 70 | 30% reduction |
| Annual sales increase | - | 2 billion yen | Exceeded the 500 million yen target |
| CO2 emission reduction | - | 500 tons per year | Environmental contribution |
Comment from management:
"PI is not merely a technological innovation; it is a transformation of the business model. We will roll it out to all plants going forward."
1.4 Conventional Method vs PI: A Workflow Comparison
As we saw in the chemical plant optimization example, conventional methods take time and money. Here, let's visually compare the workflows of the conventional method and the PI method.
Workflow Comparison Diagram
Quantitative Comparison
| Metric | Conventional method | PI method | Improvement rate |
|---|---|---|---|
| Annual number of experiments | 10-30 conditions | 100-200 conditions (made efficient by predictive support) | 5-10 times |
| Time per condition | 1-2 weeks | 1-2 weeks (experiment only) a few seconds (prediction) |
Prediction is essentially zero |
| Optimization period | 6-12 months | 2-4 months | 50-75% shorter |
| Success rate | 5-10% (rules of thumb) | 50-70% (prediction accuracy) | 5-10 times higher |
| Cost | 100 million-300 million yen | 30 million-80 million yen | 60-70% reduction |
Comparison Example on a Time Axis
When trying 50 conditions with the conventional method: - 1 condition × 1 week = 50 conditions × 1 week = 50 weeks = about 12 months
When evaluating 50 conditions with the PI method: - Data collection and model building: 2 weeks - Predicting 10,000 conditions: 1 day - Experimenting on 10 of the top 50 conditions: 10 conditions × 1 week = 10 weeks - Total: about 3 months
Time reduction: 12 months → 3 months = 75% reduction
1.5 Column: A Day in the Life of a Process Engineer
Let's see, through concrete stories, how the process development workplace has changed.
1990: The Age of Conventional Methods
A day for Engineer Tanaka (age 35)
6:00 - Arrive at the plant (early-morning inspection) Check the overnight operating status. The reactor is running under yesterday's experimental conditions.
8:00 - Sampling and analysis preparation Take a product sample from the reactor. Bring it to the analysis room.
9:00 - Quality analysis (GC, HPLC) Component analysis by gas chromatography and high-performance liquid chromatography. The measurement takes 3 hours.
12:00 - Data analysis Manually calculate peak areas and compute the yield. Enter it into Excel by hand.
14:00 - Consider the next experimental condition Looking at today's results, decide the next condition based on experience. "Let's try raising the temperature by 5°C."
16:00 - Set the reactor conditions For tomorrow's experiment, manually set the temperature controller.
18:00 - Record in the lab notebook Record today's results in detail in a handwritten lab notebook.
20:00 - Leave work
One week's output: Evaluate 1 condition, prepare the next condition
One month's output (20 days): Evaluate about 4 conditions
One year's output: Evaluate about 40-50 conditions
2025: The PI Era
A day for Engineer Sato (age 32)
9:00 - Arrive at work, check the dashboard Check on the monitor the results of the 10 conditions that the automated experimental equipment ran overnight. The data is automatically saved to a cloud database.
9:15 - Data analysis by AI The machine learning model automatically predicts yields and proposes optimizations. Analysis of the 10 conditions' data is completed in 3 minutes.
9:30 - Respond to an anomaly detection alert For 1 condition, the yield is 5% below the prediction. The root-cause analysis tool detects "an anomaly in the pressure sensor." Contact the maintenance team.
10:00 - Consider the next experimental candidates The Bayesian optimization algorithm proposes the next 20 conditions to try, based on past data and this morning's results. Predicted yields are also displayed.
10:30 - Scrutinize and select the top candidates Review the 20 proposed conditions with the engineer's eye. Considering process constraints (temperature upper limit, safe pressure range), select 10 conditions.
11:00 - Load them into the automated experimental equipment Enter the selected 10 conditions into the automated experimental system. They are scheduled to run automatically in an overnight batch.
11:30 - Team meeting Discuss this week's progress. Evaluate the performance of the new model and formulate next week's experimental plan.
13:00 - Model improvement work Add the new data obtained this week and retrain the machine learning model. Prediction accuracy improves (R²: 0.88 → 0.91).
15:00 - Report writing Review the automatically generated optimization report and create a summary for management.
17:00 - Leave work
One day's output: Evaluate 10 conditions, set the next 10 conditions for automated experiments
One month's output (20 days): Evaluate about 200 conditions
One year's output: Evaluate about 2,000 conditions
Key Points of the Change
| Item | 1990 | 2025 | Change |
|---|---|---|---|
| Evaluations per day | 0.2 conditions (1 condition per 5 days) | 10 conditions | 50 times |
| Annual evaluations | 40-50 conditions | 2,000 conditions | 40-50 times |
| Data analysis time | 3-4 hours/condition | 3 minutes/10 conditions (automated) | 99% reduction |
| Lab notebook | Handwritten | Digitized (auto-saved) | Efficient and searchable |
| Condition selection | Experience and intuition | AI proposals + human judgment | Combination |
| Report writing time | Half a day | 10 minutes (auto-generation + review) | 95% reduction |
An important point: PI is a tool that supports engineers rather than replacing them. Both Engineer Tanaka's experience and Engineer Sato's judgment are indispensable, but with the support of AI, Engineer Sato can explore more conditions, more efficiently.
1.6 Why PI "Now": Three Tailwinds
The concept of PI itself has existed since the 1990s, but it was only from the 2010s onward that it was put into practical use in earnest. Why "now"? There are three major factors.
Tailwind 1: Advances in Sensor Technology and IoT
Realizing real-time monitoring
1990s: - Number of sensors: 50-100 across the entire plant - Measurement frequency: every minute - Data storage: local servers (with capacity limits) - Cost: 500,000-1,000,000 yen per sensor
2025: - Number of sensors: 5,000-10,000 across the entire plant - Measurement frequency: every second (some every 0.1 second) - Data storage: cloud (unlimited) - Cost: 10,000-50,000 yen per sensor (90-95% reduction)
The spread of IoT (Internet of Things): - Wireless sensors: easy to install without wiring - Edge computing: preprocessing data on the sensor side - 5G communication: real-time transmission of large volumes of data
A concrete example: a smart chemical plant
A plant of a major chemical manufacturer: - Temperature sensors: 1,200 (reactors, piping, heat exchangers) - Pressure sensors: 800 - Flow rate sensors: 600 - Quality sensors: 200 (online analyzers) - Total data volume: 10 TB per day
By leveraging this data, prediction and optimization with AI became possible.
Tailwind 2: Computing Performance and the Spread of the Cloud
The benefit of Moore's Law
- 1990: Several days for one process optimization calculation
- 2000: Several hours for one calculation
- 2010: Several minutes for one calculation
- 2025: A few seconds for one calculation; 10,000 conditions in a few minutes with parallel computing
The spread of cloud computing
- Previously: In-house servers (initial investment of 100 million yen, maintenance cost of 30 million yen per year)
- Now: AWS, Google Cloud, Azure (pay-as-you-go, zero initial investment)
- Effect:
- High-performance computing is possible even for small and medium-sized enterprises
- Secure computing resources only when needed (cost reduction)
- Data sharing and collaboration are easy
Leveraging GPUs
- With the spread of deep learning, GPU computing has become common
- Machine learning models can be trained at more than 100 times the speed of a CPU
- GPU manufacturers such as NVIDIA provide industrial GPUs
Tailwind 3: Growing Social Urgency
Carbon neutrality (2050 target)
Since the 2015 Paris Agreement, chemical companies around the world have been working to reduce CO2 emissions.
CO2 emissions of Japan's chemical industry: - About 10% of all industry (about 100 million tons per year) - Target: a 46% reduction by 2030 compared to 2013
The contribution of PI: - Reducing energy consumption through process optimization (20-30%) - Reducing the energy needed for waste treatment by reducing byproducts - Accelerating the development of new green chemistry processes
A concrete example: At one petrochemical plant, optimization by PI reduced energy consumption by 25% and cut annual CO2 emissions by 10,000 tons.
Quality assurance and GMP (Good Manufacturing Practice)
In pharmaceutical manufacturing, there are strict regulations from the FDA (U.S. Food and Drug Administration) and the EMA (European Medicines Agency).
Tightening regulations: - The allowable range for batch-to-batch variation is shrinking - The introduction of Process Analytical Technology (PAT) is recommended - Real-time quality control is required
The role of PI: - Reduce batch-to-batch variation by 50-70% - Real-time quality prediction and control - Accountability to regulatory authorities (data-based decision-making)
Labor shortages and the transfer of expertise
In Japan's chemical industry, the aging of skilled technicians and the shortage of young workers are serious.
Statistics: - Average age in the chemical industry: about 45 (2023) - 30% of technicians are expected to retire over the next 10 years - Young people are turning away from manufacturing
The solution offered by PI: - Convert the tacit knowledge of skilled workers into explicit knowledge as data - Reduce the burden on young workers with AI-based operational support - Enable operation with fewer people through remote monitoring
Conclusion: PI is a technology that is needed precisely now, when technical maturity and social necessity have been satisfied simultaneously.
1.7 Chapter Summary
What We Learned
-
The history of chemical process development - From ancient distillation and fermentation to modern smart factories - Evolving from rules of thumb → scientific approach → process control → data-driven - Yet development still takes 5-10 years
-
The limitations of conventional methods - Time: 10-15 years to develop a new process - Cost: 11.6 billion-53.4 billion yen (from laboratory to commercial plant) - The difficulty of scale-up: behavior changes between the laboratory and the plant - Batch-to-batch variation: quality variability - Dependence on experience: reliance on the tacit knowledge of skilled workers
-
Lessons from chemical plant optimization - Conventional method (6 months): 24 conditions, yield +3 %pt - PI method (3 months): 20 conditions, yield +15.3 %pt - 67% shorter time, 5 times greater results
-
The advantages of PI - Annual evaluations: 40-50 conditions → 2,000 conditions (40-50 times) - Optimization period: 6-12 months → 2-4 months (50-75% shorter) - Cost reduction: 60-70% reduction
-
Why PI is needed "now" - Sensor technology and IoT (real-time data collection) - Computing performance and the cloud (fast computation, parallel processing) - Social urgency (carbon neutrality, quality assurance, labor shortages)
Important Points
- PI is a tool that supports engineers rather than replacing them
- The combination of computational prediction and experimental verification is important
- The quality and quantity of data determine prediction accuracy
- Knowledge of both chemical engineering and data science is required
On to the Next Chapter
In Chapter 2, we will study the basic concepts and methods of PI in detail: - The definition of PI and related fields - A glossary of 20 PI terms - Types of process data - The 5-step PI workflow - What are process descriptors?
Exercises
Problem 1 (Difficulty: easy)
In the history of chemical process development, explain how development methods evolved across the three eras: ancient, modern, and present day.
Hint
Think in terms of the flow from rules of thumb → scientific approach → data-driven.Sample Answer
**Ancient (BCE-around 1800)**: - Development method: rules of thumb and oral transmission - Examples: distillation technology, fermentation technology - Hundreds of years to discover optimal conditions - No scientific understanding **Modern (around 1800-1980)**: - Development method: design based on scientific principles - Example: the Haber-Bosch process (thermodynamics and reaction kinetics) - A combination of experiment and theory - Yet trial and error was still needed (Haber tested 6,500 catalysts) **Present day (1980s-present)**: - Development method: process control + data-driven - Automatic control via DCS (distributed control systems) - IoT sensors and big data - Optimization by AI/machine learning - Yet full automation has not been achieved **Conclusion**: Over thousands of years, development has evolved from rules of thumb to a scientific approach, and further to data-driven methods.Problem 2 (Difficulty: easy)
If you carry out process optimization of 50 conditions with the conventional method, how much time will it take? Assume 1 week per condition. To how many months can the PI method shorten this?
Hint
Conventional method: 50 conditions × 1 week = ? weeks PI method: 2 weeks data collection + 1 day prediction + 10 conditions × 1 week experiments = ? weeksSample Answer
**Conventional method**: - 50 conditions × 1 week = 50 weeks - 1 year = 52 weeks - **About 12 months** (almost 1 year) **PI method**: - Data collection and model building: 2 weeks - Predicting 10,000 conditions: 1 day (negligible) - Experimenting on the top 10 conditions: 10 conditions × 1 week = 10 weeks - **Total: 12 weeks = about 3 months** **Time reduction rate**: 12 months → 3 months = **75% reduction** **Conclusion**: With the PI method, the process optimization period can be shortened to one-quarter.Problem 3 (Difficulty: medium)
In the case study of chemical plant optimization, what is the biggest difference between the conventional method and the PI method? Explain from the perspectives of time, number of experiments, and results.
Hint
Using the figures in Table 1.3, compare each metric.Sample Answer
**Comparison of time**: - Conventional method: evaluated 24 conditions in 6 months, highest yield 73% (+3 %pt) - PI method: evaluated 20 conditions in 3 months, highest yield 85.3% (+15.3 %pt) - **Difference**: the PI method achieved 5 times the result in half the time **Comparison of the number of experiments**: - Conventional method: planned 50 conditions based on experience (in reality, ran out of time at 24 conditions) - PI method: efficiently selected 20 conditions via Bayesian optimization - **Difference**: the PI method achieved high results with fewer experiments **Comparison of results**: - Conventional method: yield +3 %pt, no energy reduction - PI method: yield +15.3 %pt, 30% energy reduction - **Difference**: the PI method achieved multiple goals simultaneously **The biggest difference**: The biggest difference is that **"smart exploration" using a machine learning model** became possible. With the conventional method, engineers selected conditions somewhat randomly based on experience, but with the PI method, the model learns from past data and proposes the **most promising condition to try next**. This makes it possible to reach the optimal solution with fewer experiments. Furthermore, because PI can also simultaneously evaluate **secondary effects** such as energy consumption, multi-objective optimization is possible.Data Licensing and Access
Process Databases
Industrial process data: - AIChE DIPPR Database: thermophysical property data (2,000+ substances) - URL: https://dippr.aiche.org - License: Commercial (university discounts available) - Coverage: density, viscosity, thermal conductivity, vapor pressure, etc.
- NIST Chemistry WebBook: a free thermochemical database
- URL: https://webbook.nist.gov
- License: Public domain
- Coverage: standard enthalpies of formation, reaction rate constants
Public process datasets: - Tennessee Eastman Process Dataset: an anomaly detection benchmark - URL: https://github.com/camaramm/tennessee-eastman-profBraatz - License: MIT License - Use: evaluating process control and anomaly detection algorithms
- Industrial Process Control Datasets: multiple process datasets
- URL: https://archive.ics.uci.edu/ml/datasets/Industrial+Process+Control
- License: CC BY 4.0
- Scale: 10,000-100,000 data points
Ensuring Code Reproducibility
Environment information: - Python: 3.9 or higher - Main library versions: - pandas >= 1.3.0 - numpy >= 1.21.0 - scikit-learn >= 1.0.0 - scipy >= 1.7.0
Time series analysis libraries:
- Prophet (Facebook): >= 1.1
- Installation: pip install prophet
- Use: time series forecasting of process variables
- statsmodels >= 0.13.0 (ARIMA, VAR, etc.)
- Installation:
pip install statsmodels - Use: statistical time series models
Best practices for reproducibility:
# Fix the random seed (for reproducibility of results)
import numpy as np
import random
np.random.seed(42)
random.seed(42)
# Check the Python version
import sys
print(f"Python: {sys.version}")
# Record library versions
import pandas as pd
print(f"pandas: {pd.__version__}")
Practical Pitfalls
Pitfall 1: Data Leakage in Time Series Data
Incorrect implementation example:
# ❌ Wrong: standardize on all data, then split
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X) # Learn on all data
X_train, X_test = train_test_split(X_scaled) # Split afterward
Correct implementation example:
# ✅ Correct: split first, then standardize using only the training data
X_train, X_test = train_test_split(X)
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train) # Learn on the training data only
X_test_scaled = scaler.transform(X_test) # Transform with the same parameters
Reason: Using the test data's statistics (mean, standard deviation) in training overestimates the model's performance.
Pitfall 2: Sensor Drift
Sensors accumulate measurement error over time:
# Sensor calibration check
def check_sensor_drift(data, sensor_col, window='30D'):
"""
Detect deviation from the baseline using a 30-day moving average
"""
baseline = data[sensor_col].iloc[:100].mean()
drift = data[sensor_col].rolling(window=window).mean() - baseline
if abs(drift.max()) > 0.05 * baseline:
print(f"⚠️ Sensor drift detected: {drift.max():.2f}")
return drift
Pitfall 3: Handling Missing Values
In process data, missing values occur frequently due to communication errors and sensor failures:
# Be careful with forward fill
# ❌ Wrong: unlimited forward fill
df['temperature'].fillna(method='ffill') # Fill with the previous value without limit
# ✅ Correct: set a limit on the fill
df['temperature'].fillna(method='ffill', limit=3) # Up to 3 steps at most
# Beyond that, use linear interpolation
df['temperature'].interpolate(method='linear', limit=10)
End-of-Chapter Checklist (40 items)
1. Historical Understanding (10 items)
- [ ] 1.1 Can explain the development method of ancient distillation technology (3000 BCE)
- [ ] 1.2 Can explain the scientific approach of the Haber-Bosch process (1909)
- [ ] 1.3 Remember the number of candidate catalysts Haber tested (about 6,500)
- [ ] 1.4 Can explain the impact of the introduction of DCS (1980s) on process control
- [ ] 1.5 Can list the four pillars of Industry 4.0 (2010s)
- [ ] 1.6 Can diagram the transition of the timeline of chemical process development (ancient-modern-present)
- [ ] 1.7 Can explain the evolution from rules of thumb → scientific approach → data-driven
- [ ] 1.8 Can describe the characteristics of the growth period of the petrochemical industry (1940s-1970s)
- [ ] 1.9 Can estimate the number of sensors in a smart factory (thousands)
- [ ] 1.10 Can list the four challenges revealed by the history of process development
2. Limitations of Conventional Methods (10 items)
- [ ] 2.1 Can explain the typical process development timeline (10-15 years) stage by stage
- [ ] 2.2 Can describe the breakdown of the typical cost of process development (11.6 billion-53.4 billion yen)
- [ ] 2.3 Remember the project failure rate (50-70%)
- [ ] 2.4 Can explain the four factors of the scale-up effect
- [ ] 2.5 Can calculate the scale ratio between the laboratory (100 mL) and the commercial plant (10,000 L)
- [ ] 2.6 Can analyze the cause of the scale-up failure example (yield 95% → 65%)
- [ ] 2.7 Can list the four causes of batch-to-batch variation
- [ ] 2.8 Remember the quality specification for pharmaceuticals (active ingredient 95-105%)
- [ ] 2.9 Can explain the four problems with the tacit knowledge of skilled operators
- [ ] 2.10 Can describe the significance of converting tacit knowledge into explicit knowledge with data and AI
3. Understanding the Case Study (10 items)
- [ ] 3.1 Remember the initial yield (70%) and the target (78% or higher) of the catalytic reaction process
- [ ] 3.2 Can explain the conventional method's experimental plan (50 conditions, 50 weeks)
- [ ] 3.3 Can describe the period (1 week) and scale (500 batches) of the PI team's Step 1 (data collection)
- [ ] 3.4 Can evaluate the performance of the Random Forest model (R²=0.893)
- [ ] 3.5 Remember the number of experiments (20 conditions) and the period (2 months) of Bayesian optimization
- [ ] 3.6 Can describe the final highest yield (85.3%) and the optimal conditions (temperature, pressure, etc.)
- [ ] 3.7 Remember the energy consumption reduction rate (30%) and the CO2 reduction amount (500 tons/year)
- [ ] 3.8 Can explain the development period comparison between the conventional and PI methods (12 months → 4 months, 67% shorter)
- [ ] 3.9 Remember the annual sales increase (2 billion yen, exceeding the 500 million yen target)
- [ ] 3.10 Understand the significance of the management comment ("a transformation of the business model")
4. Workflow Comparison (5 items)
- [ ] 4.1 Can explain the difference between the success rate of the conventional method (5-10%) and the PI method (50-70%)
- [ ] 4.2 Can describe the comparison of annual experiment counts (10-30 conditions vs 100-200 conditions)
- [ ] 4.3 Remember the optimization period reduction rate (50-75%)
- [ ] 4.4 Can estimate the cost reduction effect (60-70%)
- [ ] 4.5 Can diagram the difference between the two methods with a Mermaid flowchart
5. The Necessity of PI (5 items)
- [ ] 5.1 Can explain the evolution of sensor technology (50-100 in the 1990s → 5,000-10,000 in 2025)
- [ ] 5.2 Remember the sensor cost reduction rate (90-95%)
- [ ] 5.3 Can describe the evolution of computation speed (several days in 1990 → several seconds in 2025)
- [ ] 5.4 Remember the CO2 emissions of Japan's chemical industry (about 10% of all industry)
- [ ] 5.5 Understand the 2030 CO2 reduction target (46% reduction compared to 2013)
References
-
Venkatasubramanian, V. (2019). "The promise of artificial intelligence in chemical engineering: Is it here, finally?" AIChE Journal, 65(2), 466-478. DOI: 10.1002/aic.16489
-
McBride, K., & Sundmacher, K. (2019). "Overview of Surrogate Modeling in Chemical Process Engineering." Chemie Ingenieur Technik, 91(3), 228-239. DOI: 10.1002/cite.201800091
-
Lee, J. H., Shin, J., & Realff, M. J. (2018). "Machine learning: Overview of the recent progresses and implications for the process systems engineering field." Computers & Chemical Engineering, 114, 111-121. DOI: 10.1016/j.compchemeng.2017.10.008
-
The Society of Chemical Engineers, Japan (2023). "Recommendations on Promoting DX in the Chemical Process Industry." URL: https://www.scej.org
-
DIPPR Database. "AIChE Design Institute for Physical Properties." URL: https://dippr.aiche.org
-
NIST Chemistry WebBook. "NIST Standard Reference Database Number 69." URL: https://webbook.nist.gov
Author Information
Created by: MI Knowledge Hub Content Team Date created: 2025-10-16 Version: 1.1 Series: PI Introduction Series v1.0
Update history: - 2025-10-19: v1.1 Quality improvements - Added data licensing and access information (industrial DBs, public datasets) - Ensured code reproducibility (environment information, version management) - Added three practical pitfalls (data leakage, sensor drift, missing value handling) - Added 40-item end-of-chapter checklist (5 categories) - Added 2 references (DIPPR, NIST) - 2025-10-16: v1.0 First edition created - History of chemical process development (ancient-present) - Five limitations of conventional methods - Detailed case study of chemical plant optimization - Workflow comparison diagram (Mermaid) - "A Day in the Life of a Process Engineer" column (1990 vs 2025) - "Why PI Now" three tailwinds - 3 exercises
License: CC BY 4.0