Free estimate
Menu

Articles

When GPU acceleration pays off: a Monte Carlo benchmark you can rerun

Key takeaways

  • With 1,000 points or fewer, one CPU thread finished first in every sweep.
  • From 100,000 points up, the GPU beat one CPU thread in every sweep.
  • The switch came between 10,000 and 100,000 points; at 10,000 points one CPU thread was still faster in 8 of 10 sweeps.
  • The code and every timed run are public, so you can rerun the benchmark and see where the switch falls on your hardware.

Before a team ports code to a GPU, it needs to know how large a workload must be before the GPU pays for itself. This article reports one benchmark that finds that point for a simple Monte Carlo kernel, and shows how to rerun it on your own hardware.

We timed three ways to estimate pi by Monte Carlo sampling, from 1 to 1 billion random points. The three were ArrayFire on an NVIDIA GeForce RTX 4060 graphics card, one CPU thread, and 16 CPU threads with OpenMP.

Method: three ways to estimate pi

The kernel draws random points in a unit square and counts the points that fall inside the quarter circle. Four times that fraction estimates pi. Each method does the same job in its own way:

  • The GPU path uses ArrayFire to generate the points on the card and sum the hits there.
  • The one-thread path is a plain loop that draws points with the C library’s rand().
  • The 16-thread path is an OpenMP loop in which each thread has its own Mersenne Twister generator.

The CPU is an AMD Ryzen 9 5950X; under WSL2 the test had 16 of its 32 hardware threads.

Each method was timed in its own block: one untimed warm-up call, then 20 calls back to back. We report the median, and repeated the whole sweep 10 times (4 times at 1 billion points).

The timing pattern looks like this (simplified and written for this article; the full code is in the repository):

#include <algorithm>
#include <chrono>
#include <vector>

template <class F>
double median_seconds(F estimate, long points, int timed_calls = 20) {
  estimate(points);                       // one untimed warm-up call
  std::vector<double> t;
  for (int i = 0; i < timed_calls; ++i) { // then timed calls, back to back
    auto start = std::chrono::steady_clock::now();
    estimate(points);
    auto stop = std::chrono::steady_clock::now();
    t.push_back(std::chrono::duration<double>(stop - start).count());
  }
  std::nth_element(t.begin(), t.begin() + t.size() / 2, t.end());
  return t[t.size() / 2];                 // middle value of the timed calls
}

Small sizes: one CPU thread finishes first

With 1,000 points or fewer, one CPU thread finished first in every sweep.

For small problems each GPU call cost about 0.1 to 0.4 milliseconds whatever the size, while one CPU thread handled 1,000 points in about 0.01 milliseconds.

Switch point: between 10,000 and 100,000 points

From 100,000 points up, the GPU beat one CPU thread in every sweep.

Median time per pi estimate against the number of random points, on log scales, for the GPU, one CPU thread, and 16 CPU threads with OpenMP. With 1,000 points or fewer, one CPU thread finished first in every sweep. From 100,000 points up, the GPU beat one CPU thread in every sweep. The switch, marked, came between 10,000 and 100,000 points. At 1 billion points the GPU needed more memory than the card had free.1 µs1 ms1 s11,0001 million1 billionmedian time per estimaterandom pointsGPU16 CPU threadsone CPU threadthe switchmore than card’sfree memoryMedian time per pi estimate against the number of random points, on log scales, for the GPU, one CPU thread, and 16 CPU threads with OpenMP. With 1,000 points or fewer, one CPU thread finished first in every sweep. From 100,000 points up, the GPU beat one CPU thread in every sweep. The switch, marked, came between 10,000 and 100,000 points. At 1 billion points the GPU needed more memory than the card had free.1 µs1 ms1 s11,0001 million1 billionmedian time per estimaterandom pointsGPU16 CPU threadsone CPU threadthe switchmore thancard’s freememory
Fig. 1 The switch came between 10,000 and 100,000 points; at 10,000 points one CPU thread was still faster in 8 of 10 sweeps.

At 100,000 points, the 16 CPU threads still beat the GPU in every sweep.

Large sizes: the GPU pulls ahead

At 100 million points, called back to back, the GPU took about 7 milliseconds. At that size, called back to back, one CPU thread running a plain loop with the C library’s rand() took about 1 second, a ratio to the GPU of 139 to 143 in every sweep.

Against 16 CPU threads with OpenMP, the GPU was about 11 to 12 times faster at 100 million points.

One billion points: past the card’s free memory

One billion points need 8 GB on the graphics card, more than this 8 GB card had free. The GPU then took 4.0 to 4.6 seconds and lost to the 16 CPU threads, which took under 0.71 seconds.

Accuracy: every method lands near pi

At 100 million points the median estimate of every method was within 0.0001 of pi.

Test machine and kernel

These figures come from one consumer graphics card under WSL2 on Windows, measured while the GPU was busy. The test was highly parallel, with no input data to copy, and ran against a CPU loop that uses the C library’s rand() generator.

Rerun the benchmark on your hardware

The code and every timed run are public, so you can rerun the benchmark and see where the switch falls on your hardware.

The repository’s installBuildRun.sh installs ArrayFire, builds, and runs everything. To run one backend by hand after a build:

git clone --branch rerun-2026-10 https://github.com/jacobgrass/AFPiBenchmark.git
cd AFPiBenchmark
./build.sh
# device 0, 20 timed calls per size, sizes up to 10^9 points
./build/AFPiBenchMark_cuda 0 20 9

The results page on that branch lists the machine, the build flags, and the tables behind every sentence above.

Recommendations

  • Time every size your code will see, from the smallest call to the largest, because the faster method changed with size in this benchmark.
  • Compare the GPU with a multithreaded CPU baseline as well as one thread, because at 100,000 points the 16 CPU threads still beat the GPU in every sweep.
  • Check that your largest problem fits in the card’s free memory, because at 1 billion points this benchmark needed more than the card had free.
  • State the CPU loop, its random-number generator, and the call pattern next to any speed ratio, so a reader can tell what was compared.
  • Rerun the benchmark on your own hardware before a port decision; the code and every timed run are public.

Commercial teams: Industry

References

To find out whether your own code belongs on a GPU, see Scientific and HPC software or get a free estimate.

Send us the model and the cluster it has to run on.

Get a free estimate

Related case studies