CatPlat: Automated Simulation Workflows for Data-Driven Materials Discovery

Catalysis plays a critical role in modern manufacturing, driving more than 85% of its processes. In particular, heterogeneous catalysis—a chemical process where the catalyst and reactants exist in different physical states, such as solid, liquid or gas—is fundamental to petroleum refining, fertilizer production, emissions control and the conversion of renewable feedstocks into fuels and valuable chemicals. 

 

Discovering new catalysts, however, is an exceptionally complex and computationally intensive process. Researchers need to study and compare many different materials, surfaces and possible reaction sites to understand how effectively a material could act as a catalyst. This can involve running thousands of simulations before promising candidates are identified. Traditional computational workflows for heterogeneous catalysis can also involve steep learning curves, rapidly increasing complexity and time-consuming, tedious processes. 

 

To address these challenges, researchers from the Agency for Science, Technology and Research Institute of Advanced Intelligence and Computing (A*STAR IAIC) developed CatPlat (Catalysis Platform), a Python-based platform designed to automate large-scale computational simulations for catalyst research. It streamlines the workflow, from retrieving the structures of materials to setting up and executing the final calculations. 

 

On CatPlat, a user’s query is automatically converted into a series of backend tasks, which are then executed before the results are returned to the user. 

Figure 1: User interface and backend of the CatPlat platform, exemplified by CO binding process on Cu(111).

Traditional workflows for computational heterogeneous catalysis can have several limitations, such as:

 

  • High barriers to entry: Researchers often require extensive training across approximately ten different software packages before they can independently perform computational catalysis studies. This steep learning curve can make it difficult for new researchers to begin running simulations and analysing materials. 
  • Tedious and time-consuming operations: The manual preparation and submission of calculations can take several hours each day and can be prone to human error and inconsistencies.
  • Rapidly increasing complexity: As catalyst materials become more complex, particularly for low-symmetry surfaces and alloys, the number of possible sites where reactions can take place can increase significantly. Studying these possibilities can require tens of thousands of calculations. 

 

The team is also collaborating with A*STAR’s Scientific Data Strategy Office (SDSO) to develop a cloud-native version of CatPlat to meet the requirements of commercial partners. The cloud-based version uses an event-driven harness—an autonomous software framework that automatically manages and executes computational tasks—while providing more cost-efficient access to compute and high performance computing (HPC) resources. 

The CatPlat team aimed to reduce the amount of time and effort researchers spend setting up, managing and monitoring calculations, allowing them to focus on higher-level scientific questions and ultimately improving their productivity, creativity and efficiency in tackling complex scientific challenges. 

 

CatPlat enables researchers to gather data faster and more rigorously, providing deeper insights into how catalytic materials behave and how chemical reactions take place at the atomic level. The large amounts of data generated can also be used to train machine learning models to help identify promising materials for further investigation, further accelerating catalyst discovery. This deeper understanding can ultimately support the design of new and improved catalytic materials.

HPC is a key component of CatPlat, providing the computational power and scalable environment needed to run large numbers of complex simulations for materials discovery.

 

Parallel Simulations: With the ability to perform massive parallel simulations, CatPlat can run large numbers of independent simulations concurrently. This allows researchers to investigate many different materials and catalytic sites at the same time, substantially accelerating catalyst screening. 

 

Multi-node CPU and GPU resources: Some catalyst simulations involve highly detailed models containing hundreds of atoms and require substantial computational power. Multi-node CPU and GPU resources connected by high-speed interconnects enable CatPlat to perform computationally intensive simulations of large atomistic models containing more than 300 atoms in a unit cell. 

 

Together, these capabilities enable CatPlat to explore a much larger range of potential catalyst materials, often referred to as the “chemical space”, while allowing researchers to study more detailed and realistic models of how catalysts behave. HPC makes it possible to conduct these large-scale studies within timelines that would be difficult to achieve using conventional computing resources. 

CatPlat is expected to have an impact on industry and wider society over the next three to five years. 

 

Traditional materials development workflows can take 10 to 20 years from conceptualisation to deployment. As materials play an important role in many of the products and technologies used today, these long development cycles can pose a challenge when addressing urgent global issues such as climate change.

 

By accelerating the computational discovery of new catalytic materials, CatPlat can help researchers identify promising candidates within months rather than years. This allows researchers to focus subsequent testing and development on the most promising materials, potentially accelerating the development and commercial viability of greener chemical processes. 

 

Looking ahead, CatPlat could contribute to improved energy efficiency and lower greenhouse gas emissions by accelerating the discovery and design of more effective catalysts for industrial processes. It can also support emerging green technologies, such as the conversion of carbon dioxide into useful fuels and chemicals, as well as the conversion of biomass into valuable products.

 

“As Alfred North Whitehead wrote, ‘Civilisation advances by extending the number of important operations which we can perform without thinking of them.’ We envision CatPlat as that critical enabler for computational heterogeneous catalysis—harnessing the power of HPC to effortlessly scale our throughput across thousands of complex chemical simulations, transforming how we understand chemical behaviour.”

 

Dr Benjamin Chen

Senior Research Scientist

A*STAR Institute of Advanced Intelligence and Computing

Other Case Studies