This webpage hosts the materials required for the course.

Schedule

Syllabus

Unit I — Parallel Computing Platforms (8 hrs) Motivation for HPC; trends in microprocessor architectures; Flynn’s taxonomy; shared vs. distributed memory architectures; data vs. task parallelism; granularity of parallelism; interconnection networks and communication costs.

Unit II — Performance Analysis and Shared-Memory Programming (10 hrs) Sources of overhead; speedup, efficiency, and cost; Amdahl’s and Gustafson’s laws; strong/weak scaling; the roofline model; OpenMP programming model — parallel regions, work-sharing, scheduling, synchronization, and reductions.

Unit III — Distributed-Memory Programming using MPI (9 hrs) Distributed memory computing model; MPI execution model; point-to-point and collective communication; communicators; synchronization and performance considerations.

Unit IV — GPU Programming using CUDA (12 hrs) GPU architecture; CUDA programming model; kernel functions and thread hierarchy; CUDA memory hierarchy; memory coalescing and warp divergence; profiling with NVIDIA Nsight; performance optimization.

Unit V — Big Data Ecosystem using PySpark (6 hrs) Characteristics of big data (6Vs); Hadoop ecosystem overview; Apache Spark architecture; RDDs and DataFrames; transformations and actions; data processing and analytics with PySpark.

Textbooks

  • Ananth Grama, Anshul Gupta, George Karypis, Vipin Kumar, Introduction to Parallel Computing, 2nd Edition, Pearson, 2003.
  • David B. Kirk, Wen-mei W. Hwu, Programming Massively Parallel Processors: A Hands-on Approach, 4th Edition, Morgan Kaufmann, 2022.

Reference books

  • Peter S. Pacheco, Matthew Malensek, An Introduction to Parallel Programming, 2nd Edition, Morgan Kaufmann, 2021.
  • Barbara Chapman, Gabriele Jost, Ruud van der Pas, Using OpenMP: Portable Shared Memory Parallel Programming, MIT Press, 2007.
  • Jules S. Damji, Brooke Wenig, Tathagata Das, Denny Lee, Learning Spark: Lightning-Fast Data Analytics, 2nd Edition, O’Reilly, 2020.