Free estimate
Menu

Services / Scientific and HPC software

Desktop models that run on HPC clusters

We move simulation and scientific code from one workstation to batch runs on your cluster. We also design cores that teams extend, and modernize legacy code such as Fortran and C++ without losing its verification.

Get a free estimate

Proof: a desktop-only model now runs as batch jobs on a government HPC cluster

Its parameter sweeps run there as batch jobs: one scheduler job per run, many jobs per sweep. We stood up a surrogate cluster with the same scheduler and packaged releases that run without compiling on the cluster.

A desktop-only model now runs as batch jobs on an HPC cluster

Legacy code: keep, wrap, or rewrite

We keep verified Fortran and give it a C interface, so newer code in C++, Python, or another language can call it. We rewrite only when GPUs or maintenance demand it.

Is Fortran still used? Why we keep it, and how we modernize around it

GPU only where it pays

Median time per pi estimate against the number of random points, on log scales, for the GPU, one CPU thread, and 16 CPU threads with OpenMP. With 1,000 points or fewer, one CPU thread finished first in every sweep. From 100,000 points up, the GPU beat one CPU thread in every sweep. The switch, marked, came between 10,000 and 100,000 points. At 1 billion points the GPU needed more memory than the card had free.1 µs1 ms1 s11,0001 million1 billionmedian time per estimaterandom pointsGPU16 CPU threadsone CPU threadthe switchmore than card’sfree memoryMedian time per pi estimate against the number of random points, on log scales, for the GPU, one CPU thread, and 16 CPU threads with OpenMP. With 1,000 points or fewer, one CPU thread finished first in every sweep. From 100,000 points up, the GPU beat one CPU thread in every sweep. The switch, marked, came between 10,000 and 100,000 points. At 1 billion points the GPU needed more memory than the card had free.1 µs1 ms1 s11,0001 million1 billionmedian time per estimaterandom pointsGPU16 CPU threadsone CPU threadthe switchmore thancard’s freememory
Fig. 1 Against 16 CPU threads with OpenMP, the GPU was about 11 to 12 times faster at 100 million points. At 100,000 points, the 16 CPU threads still beat the GPU in every sweep.

When GPU acceleration pays off: a Monte Carlo benchmark you can rerun

Tools we work with include C++, Python, Fortran, pybind11, OpenPBS, ArrayFire, and CUDA, among others.

Case studies

We also design simulation cores that teams extend with plugins the core finds at start-up, on Windows and Linux.

Commercial teams: Industry

Can you keep our Fortran code?

Yes. We keep verified Fortran, put it on a standard build with continuous integration, and give it a C interface so newer code can call it.

Do you port code to GPUs?

Yes, and we benchmark first so the port pays. Against 16 CPU threads with OpenMP, the GPU was about 11 to 12 times faster at 100 million points. At 100,000 points, the 16 CPU threads still beat the GPU in every sweep. When GPU acceleration pays off: a Monte Carlo benchmark you can rerun

Can you work on site?

Yes, as needed and for as long as the work needs, including inside SCIFs. Our founder holds a TS/SCI clearance.

What does an estimate cost?

Estimates are free. Send the scope through the contact form; we reply within one business day and send a written quote within five business days.

Send us the model and the cluster it has to run on.

Get a free estimate