Chapter 1: Why Process Informatics?

The History and Transformation of Chemical Process Development

📖 Reading Time: 20-25 min 📊 Difficulty: Introductory 💻 Code Examples: 0 📝 Exercises: 0

Chapter 1: Why Process Informatics?

This chapter articulates the bottlenecks in process development (multiple variables, scale-up). It captures the value of data-driven optimization through a big-picture view.

💡 Supplement: In the field, designing "what and how much to measure" is essential. Starting from data visualization brings the challenges into sharp relief.

Learning Objectives

By reading this chapter, you will acquire the following: - Understand the historical evolution of chemical process development (from ancient times to the present) - Be able to explain the limitations and challenges of conventional process development - Understand the social and technical background that makes PI necessary - Learn from concrete examples of chemical plant optimization


1.1 The History of Chemical Process Development: Thousands of Years of Trial and Error

The development of the chemical industry is closely tied to the evolution of human civilization. Many of the products we use—pharmaceuticals, plastics, fuels, textiles—are produced through chemical processes.

Ancient Times: The Age of Rules of Thumb

Distillation Technology (around 3000 BCE)

Distillation, one of humanity's oldest chemical processes, was used in ancient Mesopotamia for perfume production. However, this technology was based entirely on empirical rules of thumb.

Fermentation Technology (around 7000 BCE)

The production of beer and wine is the oldest bioprocess, harnessing the metabolism of microorganisms. Although ancient people were unaware of the existence of microorganisms, they empirically discovered the appropriate conditions (temperature, sugar concentration).

The Modern Era: The Beginning of the Scientific Approach (1800s-1950s)

The Haber-Bosch Process (1909)

German chemists Fritz Haber and Carl Bosch developed an industrial process to synthesize ammonia from nitrogen and hydrogen. This was the beginning of design based on scientific principles.

However, discovering the optimal catalyst (iron-based) still required an enormous amount of trial and error. Haber is said to have tested about 6,500 candidate catalysts.

The Development of the Petrochemical Industry (1940s-1970s)

After World War II, the chemical industry using petroleum as a feedstock developed rapidly. From basic chemicals such as ethylene and propylene, plastics such as polyethylene and polypropylene came to be mass-produced.

The Present Day: The Age of Process Control (1980s-Present)

Introduction of DCS (Distributed Control System) (1980s)

Distributed control systems made it possible to automatically control the temperature, pressure, and flow rate of the entire plant. This significantly reduced the burden on human operators.

Smart Factories and Industry 4.0 (2010s-Present)

Through the integration of IoT sensors, big data, and AI/machine learning, real-time optimization is becoming possible.

Challenges Revealed by History

Looking back over thousands of years of chemical process development history, the following challenges emerge:

  1. It takes time: 5-10 years to develop a new process, and several more years for scale-up
  2. It is expensive: Hundreds of millions of yen to build a pilot plant, tens of billions of yen to build a full-scale plant
  3. Dependence on trial and error: An enormous number of experiments are needed to discover optimal conditions
  4. Dependence on tacit knowledge: The experience and intuition of skilled operators are indispensable

A question: What if we could systematically optimize chemical processes with data and AI?

This is the starting point of Process Informatics (PI), the next-generation process development method.


1.2 The Limitations of Conventional Process Development

Modern chemical process development has become far more sophisticated than in ancient times. Nevertheless, it still faces major challenges.

Challenge 1: It Takes Time

A typical process development timeline

Years 1-2: Literature review and laboratory-scale experiments
  ↓
Years 3-5: Bench-scale experiments (100 mL-1 L scale)
  ↓
Years 6-8: Pilot plant construction and operation (10-100 L scale)
  ↓
Years 9-12: Scale-up optimization
  ↓
Years 13-15: Commercial plant construction and trial operation (1,000-10,000 L scale)
  ↓
Years 16-20: Start of commercial production and continuous optimization

Result: Commercializing a new process takes 10-15 years on average [1,2].

Challenge 2: It Is Expensive

Typical costs of process development (medium-scale chemical plant)

Development stage Cost (yen) Period
Laboratory research 50 million-100 million yen 2-3 years
Bench scale 100 million-300 million yen 2-3 years
Pilot plant 1 billion-3 billion yen 3-5 years
Commercial plant 10 billion-50 billion yen 3-5 years
Total 11.6 billion-53.4 billion yen 10-16 years

Because 50-70% of projects never reach commercialization, the cost of failure is also enormous.

Challenge 3: The Difficulty of Scale-Up

Problems in scale-up

Chemical reactions behave differently depending on the size of the reactor. This is called the "scale-up effect."

Differences between the laboratory (100 mL) and the commercial plant (10,000 L):

Factor Laboratory Commercial plant Effect
Mixing efficiency Uniform Non-uniform (dead zones exist) Reduced reaction rate
Heat removal Easy Difficult (small surface-area-to-volume ratio) Temperature rise, increased side reactions
Residence time distribution Uniform Non-uniform Reduced yield
Scale ratio 1 100,000 times Hydrodynamic behavior changes

A concrete example: a failed scale-up of a pharmaceutical process

Cause: Because it was an exothermic reaction, heat removal could not keep up at large scale, and side reactions increased.

Time required to address it: Two years of additional experiments and 300 million yen of additional investment.

Challenge 4: Batch-to-Batch Variation

The problem of quality variability

In chemical processes, even when you intend to operate under the same conditions, product quality varies from batch to batch (production lot).

Causes of variation: - Variation in raw materials (purity, impurities) - Changes in environmental conditions (outside temperature, humidity) - Aging degradation of equipment - Differences in operator skill

A concrete example: pharmaceutical manufacturing

Pharmaceuticals require strict quality specifications. Under the standards of the FDA (U.S. Food and Drug Administration): - Active ingredient content: within the range of 95-105% - Impurities: <0.1%

Conventional management methods: - Record the measurement results of each batch - Discard batches that fall outside specifications (cost loss) - Weeks spent investigating the cause

Problem: Reactive, after-the-fact responses cannot prevent quality defects.

Challenge 5: Dependence on Experience and Intuition

The tacit knowledge of skilled operators

In chemical plants, the "intuition" of skilled operators is important:

Such tacit knowledge is extremely valuable, but it has the following problems:

  1. Difficult to systematize: Because it is based on personal experience, it is hard to put into words
  2. Reproducibility problem: Even under the same conditions, judgments differ from person to person
  3. Time to train young workers: Becoming skilled requires 10-20 years of experience
  4. Loss of knowledge through retirement: When veterans retire, the expertise is lost

A question: What if we could convert tacit knowledge into explicit knowledge with data and AI?


1.3 Case Study: A Real Example of Chemical Plant Optimization

Let's take a detailed look at a catalytic reaction process optimization project at a major chemical manufacturer. This is a typical example illustrating the difficulty of conventional methods and the potential of PI.

Background: The Challenge of Improving Yield

Product: A fine chemical product (a pharmaceutical intermediate) Process: A catalytic reaction (three reaction stages proceed within the reactor) Challenge: The yield had stalled at 70%, while the industry leader achieved 85%

The demand from management:

"Improve the yield to 78% or higher and increase annual sales by 500 million yen. The deadline is 12 months."

Phase 1: Optimization by Conventional Methods (6 Months)

Approach: Trial and error by experienced engineers

Variable parameters: - Reaction temperature (80-120°C) - Pressure (1-5 atm) - Catalyst amount (1-5 wt%) - Residence time (30-180 minutes) - Feed flow rate (10-50 L/h)

Experimental design: - Each parameter varied at 3 levels - Total combinations: 3^5 = 243 - In practice, 50 conditions were selected based on experience

Experimental schedule: - 1 week per condition (preparation, reaction, analysis) - 50 conditions × 1 week = 50 weeks (about 12 months)

Result (at the 6-month mark): - Number of experiments completed: 24 conditions - Highest yield: 73% (+3 %pt) - Would not meet the deadline

Phase 2: Introduction of the PI Method (3 Months)

By management's decision, the project was handed over to the PI team.

Step 1: Collect existing data (1 week)

Operating data from the past 5 years was collected: - Process parameters: temperature, pressure, flow rate, catalyst amount (every second) - Product quality data: yield, selectivity, impurities (per batch) - Total amount of data: about 500 batches, 1,000,000 data points each

Step 2: Build a machine learning model (2 weeks)

Model used: Random Forest

# Simplified code example
from sklearn.ensemble import RandomForestRegressor

# Descriptors (inputs): temperature, pressure, catalyst amount, residence time, flow rate
X = process_data[['temperature', 'pressure', 'catalyst', 'residence_time', 'flow_rate']]

# Target (output): yield
y = process_data['yield']

# Model training
model = RandomForestRegressor(n_estimators=100)
model.fit(X, y)

# Prediction accuracy
print(f"R² score: {model.score(X_test, y_test):.3f}")
# Output: R² score: 0.893 (high accuracy)

Model performance: - R² = 0.893 (89.3% prediction accuracy) - MAE (mean absolute error) = 2.1% (yield prediction error of ±2.1%)

Step 3: Exploration by Bayesian optimization (2 months)

Using the machine learning model, optimal conditions were efficiently explored.

Number of conditions explored: 20 conditions (the conventional plan called for 50 conditions)

How Bayesian optimization works: 1. The model proposes the "next condition to try" 2. Verify it with an experiment 3. Add the result to the data and update the model 4. Repeat

Result (at the 3-month mark): - Highest yield: 85.3% (+15.3 %pt, exceeding the 78% target) - Optimal conditions: - Temperature: 102°C (previously 95°C) - Pressure: 3.2 atm (previously 4.0 atm) - Catalyst amount: 2.8 wt% (previously 3.5 wt%) - Residence time: 65 minutes (previously 90 minutes)

Phase 3: Verification at the Actual Plant (1 Month)

Pilot test: Confirmed the predicted yield of 85%

Rollout to the commercial plant: - First batch: 84.2% yield - After stable operation (10 batches): average yield 85.1%

Additional findings: - Energy consumption reduced by 30% (the effect of lower temperature and shorter time) - Byproducts reduced by 40% (improved selectivity)

Results and Impact

Metric Conventional After PI Improvement rate
Yield 70% 85.3% +15.3 %pt
Development period 12 months (planned) 4 months (actual) 67% shorter
Number of experiments 50 (planned) 20 (actual) 60% fewer
Energy consumption 100 (baseline) 70 30% reduction
Annual sales increase - 2 billion yen Exceeded the 500 million yen target
CO2 emission reduction - 500 tons per year Environmental contribution

Comment from management:

"PI is not merely a technological innovation; it is a transformation of the business model. We will roll it out to all plants going forward."


1.4 Conventional Method vs PI: A Workflow Comparison

As we saw in the chemical plant optimization example, conventional methods take time and money. Here, let's visually compare the workflows of the conventional method and the PI method.

Workflow Comparison Diagram

flowchart TD subgraph "Conventional Method (Trial-and-Error Type)" A1[Condition selection based on experience] -->|1 week| A2[Conduct experiment 1] A2 -->|1 week| A3[Data analysis] A3 -->|1 week| A4{Goal achieved?} A4 -->|No 95%| A1 A4 -->|Yes 5%| A5[Deploy to plant] style A1 fill:#ffcccc style A2 fill:#ffcccc style A3 fill:#ffcccc style A4 fill:#ffcccc style A5 fill:#ccffcc end subgraph "PI Method (Data-Driven Type)" B1[Collect past data] -->|1 week| B2[Build machine learning model] B2 -->|1 week| B3[Predict 10,000 conditions] B3 -->|1 day| B4[Narrow to top 100 conditions] B4 -->|1 day| B5[Select top 10 conditions] B5 -->|2 months| B6[Experimentally verify 10 conditions] B6 -->|1 week| B7{Goal achieved?} B7 -->|No 30%| B8[Add data and retrain] B8 -->|1 week| B2 B7 -->|Yes 70%| B9[Deploy to plant] style B1 fill:#ccddff style B2 fill:#ccddff style B3 fill:#ccddff style B4 fill:#ccddff style B5 fill:#ccddff style B6 fill:#ffffcc style B7 fill:#ccddff style B8 fill:#ccddff style B9 fill:#ccffcc end A1 -.compare.- B1

Quantitative Comparison

Metric Conventional method PI method Improvement rate
Annual number of experiments 10-30 conditions 100-200 conditions (made efficient by predictive support) 5-10 times
Time per condition 1-2 weeks 1-2 weeks (experiment only)
a few seconds (prediction)
Prediction is essentially zero
Optimization period 6-12 months 2-4 months 50-75% shorter
Success rate 5-10% (rules of thumb) 50-70% (prediction accuracy) 5-10 times higher
Cost 100 million-300 million yen 30 million-80 million yen 60-70% reduction

Comparison Example on a Time Axis

When trying 50 conditions with the conventional method: - 1 condition × 1 week = 50 conditions × 1 week = 50 weeks = about 12 months

When evaluating 50 conditions with the PI method: - Data collection and model building: 2 weeks - Predicting 10,000 conditions: 1 day - Experimenting on 10 of the top 50 conditions: 10 conditions × 1 week = 10 weeks - Total: about 3 months

Time reduction: 12 months → 3 months = 75% reduction


1.5 Column: A Day in the Life of a Process Engineer

Let's see, through concrete stories, how the process development workplace has changed.

1990: The Age of Conventional Methods

A day for Engineer Tanaka (age 35)

6:00 - Arrive at the plant (early-morning inspection) Check the overnight operating status. The reactor is running under yesterday's experimental conditions.

8:00 - Sampling and analysis preparation Take a product sample from the reactor. Bring it to the analysis room.

9:00 - Quality analysis (GC, HPLC) Component analysis by gas chromatography and high-performance liquid chromatography. The measurement takes 3 hours.

12:00 - Data analysis Manually calculate peak areas and compute the yield. Enter it into Excel by hand.

14:00 - Consider the next experimental condition Looking at today's results, decide the next condition based on experience. "Let's try raising the temperature by 5°C."

16:00 - Set the reactor conditions For tomorrow's experiment, manually set the temperature controller.

18:00 - Record in the lab notebook Record today's results in detail in a handwritten lab notebook.

20:00 - Leave work

One week's output: Evaluate 1 condition, prepare the next condition

One month's output (20 days): Evaluate about 4 conditions

One year's output: Evaluate about 40-50 conditions

2025: The PI Era

A day for Engineer Sato (age 32)

9:00 - Arrive at work, check the dashboard Check on the monitor the results of the 10 conditions that the automated experimental equipment ran overnight. The data is automatically saved to a cloud database.

9:15 - Data analysis by AI The machine learning model automatically predicts yields and proposes optimizations. Analysis of the 10 conditions' data is completed in 3 minutes.

9:30 - Respond to an anomaly detection alert For 1 condition, the yield is 5% below the prediction. The root-cause analysis tool detects "an anomaly in the pressure sensor." Contact the maintenance team.

10:00 - Consider the next experimental candidates The Bayesian optimization algorithm proposes the next 20 conditions to try, based on past data and this morning's results. Predicted yields are also displayed.

10:30 - Scrutinize and select the top candidates Review the 20 proposed conditions with the engineer's eye. Considering process constraints (temperature upper limit, safe pressure range), select 10 conditions.

11:00 - Load them into the automated experimental equipment Enter the selected 10 conditions into the automated experimental system. They are scheduled to run automatically in an overnight batch.

11:30 - Team meeting Discuss this week's progress. Evaluate the performance of the new model and formulate next week's experimental plan.

13:00 - Model improvement work Add the new data obtained this week and retrain the machine learning model. Prediction accuracy improves (R²: 0.88 → 0.91).

15:00 - Report writing Review the automatically generated optimization report and create a summary for management.

17:00 - Leave work

One day's output: Evaluate 10 conditions, set the next 10 conditions for automated experiments

One month's output (20 days): Evaluate about 200 conditions

One year's output: Evaluate about 2,000 conditions

Key Points of the Change

Item 1990 2025 Change
Evaluations per day 0.2 conditions (1 condition per 5 days) 10 conditions 50 times
Annual evaluations 40-50 conditions 2,000 conditions 40-50 times
Data analysis time 3-4 hours/condition 3 minutes/10 conditions (automated) 99% reduction
Lab notebook Handwritten Digitized (auto-saved) Efficient and searchable
Condition selection Experience and intuition AI proposals + human judgment Combination
Report writing time Half a day 10 minutes (auto-generation + review) 95% reduction

An important point: PI is a tool that supports engineers rather than replacing them. Both Engineer Tanaka's experience and Engineer Sato's judgment are indispensable, but with the support of AI, Engineer Sato can explore more conditions, more efficiently.


1.6 Why PI "Now": Three Tailwinds

The concept of PI itself has existed since the 1990s, but it was only from the 2010s onward that it was put into practical use in earnest. Why "now"? There are three major factors.

Tailwind 1: Advances in Sensor Technology and IoT

Realizing real-time monitoring

1990s: - Number of sensors: 50-100 across the entire plant - Measurement frequency: every minute - Data storage: local servers (with capacity limits) - Cost: 500,000-1,000,000 yen per sensor

2025: - Number of sensors: 5,000-10,000 across the entire plant - Measurement frequency: every second (some every 0.1 second) - Data storage: cloud (unlimited) - Cost: 10,000-50,000 yen per sensor (90-95% reduction)

The spread of IoT (Internet of Things): - Wireless sensors: easy to install without wiring - Edge computing: preprocessing data on the sensor side - 5G communication: real-time transmission of large volumes of data

A concrete example: a smart chemical plant

A plant of a major chemical manufacturer: - Temperature sensors: 1,200 (reactors, piping, heat exchangers) - Pressure sensors: 800 - Flow rate sensors: 600 - Quality sensors: 200 (online analyzers) - Total data volume: 10 TB per day

By leveraging this data, prediction and optimization with AI became possible.

Tailwind 2: Computing Performance and the Spread of the Cloud

The benefit of Moore's Law

The spread of cloud computing

Leveraging GPUs

Tailwind 3: Growing Social Urgency

Carbon neutrality (2050 target)

Since the 2015 Paris Agreement, chemical companies around the world have been working to reduce CO2 emissions.

CO2 emissions of Japan's chemical industry: - About 10% of all industry (about 100 million tons per year) - Target: a 46% reduction by 2030 compared to 2013

The contribution of PI: - Reducing energy consumption through process optimization (20-30%) - Reducing the energy needed for waste treatment by reducing byproducts - Accelerating the development of new green chemistry processes

A concrete example: At one petrochemical plant, optimization by PI reduced energy consumption by 25% and cut annual CO2 emissions by 10,000 tons.

Quality assurance and GMP (Good Manufacturing Practice)

In pharmaceutical manufacturing, there are strict regulations from the FDA (U.S. Food and Drug Administration) and the EMA (European Medicines Agency).

Tightening regulations: - The allowable range for batch-to-batch variation is shrinking - The introduction of Process Analytical Technology (PAT) is recommended - Real-time quality control is required

The role of PI: - Reduce batch-to-batch variation by 50-70% - Real-time quality prediction and control - Accountability to regulatory authorities (data-based decision-making)

Labor shortages and the transfer of expertise

In Japan's chemical industry, the aging of skilled technicians and the shortage of young workers are serious.

Statistics: - Average age in the chemical industry: about 45 (2023) - 30% of technicians are expected to retire over the next 10 years - Young people are turning away from manufacturing

The solution offered by PI: - Convert the tacit knowledge of skilled workers into explicit knowledge as data - Reduce the burden on young workers with AI-based operational support - Enable operation with fewer people through remote monitoring

Conclusion: PI is a technology that is needed precisely now, when technical maturity and social necessity have been satisfied simultaneously.


1.7 Chapter Summary

What We Learned

  1. The history of chemical process development - From ancient distillation and fermentation to modern smart factories - Evolving from rules of thumb → scientific approach → process control → data-driven - Yet development still takes 5-10 years

  2. The limitations of conventional methods - Time: 10-15 years to develop a new process - Cost: 11.6 billion-53.4 billion yen (from laboratory to commercial plant) - The difficulty of scale-up: behavior changes between the laboratory and the plant - Batch-to-batch variation: quality variability - Dependence on experience: reliance on the tacit knowledge of skilled workers

  3. Lessons from chemical plant optimization - Conventional method (6 months): 24 conditions, yield +3 %pt - PI method (3 months): 20 conditions, yield +15.3 %pt - 67% shorter time, 5 times greater results

  4. The advantages of PI - Annual evaluations: 40-50 conditions → 2,000 conditions (40-50 times) - Optimization period: 6-12 months → 2-4 months (50-75% shorter) - Cost reduction: 60-70% reduction

  5. Why PI is needed "now" - Sensor technology and IoT (real-time data collection) - Computing performance and the cloud (fast computation, parallel processing) - Social urgency (carbon neutrality, quality assurance, labor shortages)

Important Points

On to the Next Chapter

In Chapter 2, we will study the basic concepts and methods of PI in detail: - The definition of PI and related fields - A glossary of 20 PI terms - Types of process data - The 5-step PI workflow - What are process descriptors?


Exercises

Problem 1 (Difficulty: easy)

In the history of chemical process development, explain how development methods evolved across the three eras: ancient, modern, and present day.

Hint Think in terms of the flow from rules of thumb → scientific approach → data-driven.
Sample Answer **Ancient (BCE-around 1800)**: - Development method: rules of thumb and oral transmission - Examples: distillation technology, fermentation technology - Hundreds of years to discover optimal conditions - No scientific understanding **Modern (around 1800-1980)**: - Development method: design based on scientific principles - Example: the Haber-Bosch process (thermodynamics and reaction kinetics) - A combination of experiment and theory - Yet trial and error was still needed (Haber tested 6,500 catalysts) **Present day (1980s-present)**: - Development method: process control + data-driven - Automatic control via DCS (distributed control systems) - IoT sensors and big data - Optimization by AI/machine learning - Yet full automation has not been achieved **Conclusion**: Over thousands of years, development has evolved from rules of thumb to a scientific approach, and further to data-driven methods.

Problem 2 (Difficulty: easy)

If you carry out process optimization of 50 conditions with the conventional method, how much time will it take? Assume 1 week per condition. To how many months can the PI method shorten this?

Hint Conventional method: 50 conditions × 1 week = ? weeks PI method: 2 weeks data collection + 1 day prediction + 10 conditions × 1 week experiments = ? weeks
Sample Answer **Conventional method**: - 50 conditions × 1 week = 50 weeks - 1 year = 52 weeks - **About 12 months** (almost 1 year) **PI method**: - Data collection and model building: 2 weeks - Predicting 10,000 conditions: 1 day (negligible) - Experimenting on the top 10 conditions: 10 conditions × 1 week = 10 weeks - **Total: 12 weeks = about 3 months** **Time reduction rate**: 12 months → 3 months = **75% reduction** **Conclusion**: With the PI method, the process optimization period can be shortened to one-quarter.

Problem 3 (Difficulty: medium)

In the case study of chemical plant optimization, what is the biggest difference between the conventional method and the PI method? Explain from the perspectives of time, number of experiments, and results.

Hint Using the figures in Table 1.3, compare each metric.
Sample Answer **Comparison of time**: - Conventional method: evaluated 24 conditions in 6 months, highest yield 73% (+3 %pt) - PI method: evaluated 20 conditions in 3 months, highest yield 85.3% (+15.3 %pt) - **Difference**: the PI method achieved 5 times the result in half the time **Comparison of the number of experiments**: - Conventional method: planned 50 conditions based on experience (in reality, ran out of time at 24 conditions) - PI method: efficiently selected 20 conditions via Bayesian optimization - **Difference**: the PI method achieved high results with fewer experiments **Comparison of results**: - Conventional method: yield +3 %pt, no energy reduction - PI method: yield +15.3 %pt, 30% energy reduction - **Difference**: the PI method achieved multiple goals simultaneously **The biggest difference**: The biggest difference is that **"smart exploration" using a machine learning model** became possible. With the conventional method, engineers selected conditions somewhat randomly based on experience, but with the PI method, the model learns from past data and proposes the **most promising condition to try next**. This makes it possible to reach the optimal solution with fewer experiments. Furthermore, because PI can also simultaneously evaluate **secondary effects** such as energy consumption, multi-objective optimization is possible.

Data Licensing and Access

Process Databases

Industrial process data: - AIChE DIPPR Database: thermophysical property data (2,000+ substances) - URL: https://dippr.aiche.org - License: Commercial (university discounts available) - Coverage: density, viscosity, thermal conductivity, vapor pressure, etc.

Public process datasets: - Tennessee Eastman Process Dataset: an anomaly detection benchmark - URL: https://github.com/camaramm/tennessee-eastman-profBraatz - License: MIT License - Use: evaluating process control and anomaly detection algorithms

Ensuring Code Reproducibility

Environment information: - Python: 3.9 or higher - Main library versions: - pandas >= 1.3.0 - numpy >= 1.21.0 - scikit-learn >= 1.0.0 - scipy >= 1.7.0

Time series analysis libraries: - Prophet (Facebook): >= 1.1 - Installation: pip install prophet - Use: time series forecasting of process variables

Best practices for reproducibility:

# Fix the random seed (for reproducibility of results)
import numpy as np
import random
np.random.seed(42)
random.seed(42)

# Check the Python version
import sys
print(f"Python: {sys.version}")

# Record library versions
import pandas as pd
print(f"pandas: {pd.__version__}")

Practical Pitfalls

Pitfall 1: Data Leakage in Time Series Data

Incorrect implementation example:

# ❌ Wrong: standardize on all data, then split
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)  # Learn on all data
X_train, X_test = train_test_split(X_scaled)  # Split afterward

Correct implementation example:

# ✅ Correct: split first, then standardize using only the training data
X_train, X_test = train_test_split(X)
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)  # Learn on the training data only
X_test_scaled = scaler.transform(X_test)  # Transform with the same parameters

Reason: Using the test data's statistics (mean, standard deviation) in training overestimates the model's performance.

Pitfall 2: Sensor Drift

Sensors accumulate measurement error over time:

# Sensor calibration check
def check_sensor_drift(data, sensor_col, window='30D'):
    """
    Detect deviation from the baseline using a 30-day moving average
    """
    baseline = data[sensor_col].iloc[:100].mean()
    drift = data[sensor_col].rolling(window=window).mean() - baseline
    if abs(drift.max()) > 0.05 * baseline:
        print(f"⚠️ Sensor drift detected: {drift.max():.2f}")
    return drift

Pitfall 3: Handling Missing Values

In process data, missing values occur frequently due to communication errors and sensor failures:

# Be careful with forward fill
# ❌ Wrong: unlimited forward fill
df['temperature'].fillna(method='ffill')  # Fill with the previous value without limit

# ✅ Correct: set a limit on the fill
df['temperature'].fillna(method='ffill', limit=3)  # Up to 3 steps at most
# Beyond that, use linear interpolation
df['temperature'].interpolate(method='linear', limit=10)

End-of-Chapter Checklist (40 items)

1. Historical Understanding (10 items)

2. Limitations of Conventional Methods (10 items)

3. Understanding the Case Study (10 items)

4. Workflow Comparison (5 items)

5. The Necessity of PI (5 items)


References

  1. Venkatasubramanian, V. (2019). "The promise of artificial intelligence in chemical engineering: Is it here, finally?" AIChE Journal, 65(2), 466-478. DOI: 10.1002/aic.16489

  2. McBride, K., & Sundmacher, K. (2019). "Overview of Surrogate Modeling in Chemical Process Engineering." Chemie Ingenieur Technik, 91(3), 228-239. DOI: 10.1002/cite.201800091

  3. Lee, J. H., Shin, J., & Realff, M. J. (2018). "Machine learning: Overview of the recent progresses and implications for the process systems engineering field." Computers & Chemical Engineering, 114, 111-121. DOI: 10.1016/j.compchemeng.2017.10.008

  4. The Society of Chemical Engineers, Japan (2023). "Recommendations on Promoting DX in the Chemical Process Industry." URL: https://www.scej.org

  5. DIPPR Database. "AIChE Design Institute for Physical Properties." URL: https://dippr.aiche.org

  6. NIST Chemistry WebBook. "NIST Standard Reference Database Number 69." URL: https://webbook.nist.gov


Author Information

Created by: MI Knowledge Hub Content Team Date created: 2025-10-16 Version: 1.1 Series: PI Introduction Series v1.0

Update history: - 2025-10-19: v1.1 Quality improvements - Added data licensing and access information (industrial DBs, public datasets) - Ensured code reproducibility (environment information, version management) - Added three practical pitfalls (data leakage, sensor drift, missing value handling) - Added 40-item end-of-chapter checklist (5 categories) - Added 2 references (DIPPR, NIST) - 2025-10-16: v1.0 First edition created - History of chemical process development (ancient-present) - Five limitations of conventional methods - Detailed case study of chemical plant optimization - Workflow comparison diagram (Mermaid) - "A Day in the Life of a Process Engineer" column (1990 vs 2025) - "Why PI Now" three tailwinds - 3 exercises

License: CC BY 4.0

Disclaimer