At COMPAS 2026, Thomas Hérault highlighted the critical role of fault tolerance in the future of High-Performance Computing, from theoretical foundations to practical solutions for next-generation HPC systems.

From 30 June to 3 July 2026, COMPAS 2026, the French-speaking conference on parallelism, architecture and systems, brought together researchers, young scientists and industry representatives in Anglet, France. Organised by the Université de Pau et des Pays de l’Adour (UPPA) and the LIUPPA laboratory, the conference provided a forum for discussing the latest developments in parallel computing, computer architectures and systems.

The 2026 edition featured a rich scientific programme covering topics ranging from high-performance computing and hardware architectures to distributed systems, cloud computing, energy efficiency, artificial intelligence and resilience. With 128 participants, 67 presentations and 20 posters, COMPAS 2026 highlighted the vitality of the French-speaking research community and fostered exchanges between established researchers, early-career scientists and industry.

Thomas Hérault: addressing fault tolerance at scale

One of the conference’s three invited lectures was delivered by Thomas Hérault (Inria Bordeaux). His keynote, “Fault-Tolerance in High Performance Computing: from Theory to Practice”, addressed one of the key challenges facing the future of large-scale computing: how to design HPC systems and software that can continue operating efficiently despite increasingly frequent hardware and software failures.

As computing systems grow towards Exascale, the sheer number of components involved makes resilience and fault tolerance increasingly important. Thomas Hérault’s presentation explored this challenge from fundamental theoretical concepts through to their implementation in modern HPC environments, including next-generation MPI systems.

His contribution also complemented a dedicated Fault-tolerance session, which he chaired during the conference. The session brought together research on fault tolerance for task-based runtime systems, fault detection in programmable switches and distributed-system resilience, illustrating the breadth of approaches being developed to make future HPC infrastructures more robust.

Why it matters for NumPEx

Resilience is a fundamental challenge for Exascale computing. As systems become larger and more heterogeneous, the probability of failures during long-running scientific simulations increases, making fault tolerance an essential component of the Exascale software stack.

Through its support of COMPAS 2026 and the participation of researchers such as Thomas Hérault, NumPEx contributes to strengthening the French HPC community and to promoting exchanges around the software challenges that will shape the next generation of computing systems.

By connecting research on algorithms, runtime systems, architectures and applications, events such as COMPAS provide an important forum for anticipating these challenges and developing the technologies required to make Exascale computing reliable, efficient and usable for scientific applications.


NumPEx Newsletter

Subscribe to our newsletter to stay informed on the latest breakthroughs in High-Performance Computing, Exascale research, and cutting-edge digital innovations.

Privacy Preference Center