Parameterized simulation runs on a PBS cluster from one spreadsheet
Key takeaways
- Stand up a surrogate cluster with the same scheduler first, so every script is proven before it reaches the real cluster.
- Package the model so it runs on the cluster without compiling anything there.
- Turn each row of a design-of-experiments spreadsheet into a run directory and a scheduler job with its own logs.
- When the cluster has no shared filesystem, run the whole batch on one large node.
- Ship the post-run analysis inside the release, so the results can be read with no network.
Case study: A desktop-only model now runs as batch jobs on an HPC cluster
A simulation that runs on one desktop answers one question at a time. A study needs many variations: different parameters, different start times, different settings for each subsystem. A cluster with a batch scheduler can run them side by side, if the model and its inputs are ready for it.
This guide shows how we turned a desktop-only model into batches of PBS jobs on a customer’s air-gapped cluster, starting from one spreadsheet. The customer side is in the case study A desktop-only model now runs as batch jobs on an HPC cluster.
Build a surrogate cluster with the same scheduler
Time on the real cluster is scarce, and it sits on a network you cannot reach from your desk. We stood up a surrogate cluster with the same scheduler, so that every job script and every package ran there first.
The surrogate later became a container stack whose worker count is a launch parameter. Each worker registers itself with the scheduler’s server when it starts, so a surrogate with two workers or twenty comes from the same command:
# Illustrative compose file: one PBS server and any number of workers
services:
pbs-server:
image: surrogate/openpbs:latest
hostname: pbs-server
command: ["/opt/surrogate/start-server.sh"]
worker:
image: surrogate/openpbs:latest
depends_on: [pbs-server]
# registers this container as a node, then runs the execution daemon
command: ["/opt/surrogate/start-worker.sh"]
docker compose up -d --scale worker=4
Match the parts of the cluster that jobs can see: the scheduler and its version, the operating system of the compute nodes, the queue names, and the resource limits. Hardware does not need to match. A surrogate proves that scripts, paths, and packages are right, not how fast the real cluster runs.
Keep the surrogate’s definition in version control beside the job scripts. A new developer can then start the same cluster on a laptop and test a change before it goes anywhere near the customer’s network.
Package the model so nothing compiles on the cluster
A release that must be compiled on the cluster depends on whatever compilers the cluster happens to have. We packaged releases that run without compiling on the cluster: the model, its libraries, and its runtime travel in one archive that unpacks into a single folder.
Test the package the way a job sees it: unpacked on a compute node of the surrogate, started by the scheduler, with a clean environment. A job inherits far less than an interactive shell does, so a package that works at a login prompt can still miss a library inside a job:
# Inside a job script: start the model with a clean, explicit environment
RELEASE=/path/to/release # where the release archive was unpacked
env -i HOME="$HOME" PATH=/usr/bin:/bin \
"$RELEASE/bin/model" --version
For a cluster whose operating system is older than the build machine’s, the method is in Running modern software on an operating system past end of maintenance.
Turn spreadsheet rows into run directories
Analysts already plan studies as a design-of-experiments spreadsheet: one row per run, one column per parameter. We wrote the tooling that turns that spreadsheet into a batch of parameterized runs. The spreadsheet stays the single place where a study is defined.
Each row becomes a run directory holding the model’s input file, a copy of its parameters, and, later, its logs and results. A sketch of the generator, written for this article:
import csv, json
from pathlib import Path
from string import Template
template = Template(Path("templates/input.tmpl").read_text())
runs = Path("runs")
with open("study.csv", newline="") as f:
for i, row in enumerate(csv.DictReader(f), start=1):
run = runs / f"run-{i:04d}"
run.mkdir(parents=True, exist_ok=True)
(run / "params.json").write_text(json.dumps(row, indent=2))
(run / "input.txt").write_text(template.substitute(row))
Template.substitute raises an error when the template names a column the spreadsheet lacks. That makes a typo in a column name stop the generator, before any job reaches the queue.
Submit one scheduler job per run, with its own logs
Each run becomes its own scheduler job with its own logs. A run that stops early then shows up as one directory with its own error log, and the rest of the batch carries on.
There are two ways to submit the batch. One job script per run is the simplest to read and to rerun. A PBS job array submits the whole batch at once, and each subjob picks its row from its index:
#!/bin/bash
#PBS -N study
#PBS -J 1-100 # illustrative: one subjob per spreadsheet row
#PBS -l select=1:ncpus=1
#PBS -l walltime=04:00:00
#PBS -v RELEASE=/path/to/release
RUN=$(printf "run-%04d" "$PBS_ARRAY_INDEX")
cd "$PBS_O_WORKDIR/runs/$RUN" || exit 1
exec > stdout.log 2> stderr.log
echo running > status
if "$RELEASE/bin/model" input.txt; then echo done > status
else echo "exit $?" > status; fi
The status file in each run directory gives a quick view of the batch with one command, cat runs/*/status | sort | uniq -c, without asking the scheduler.
Request resources the way the model uses them
Ask the scheduler for what one run needs, not for what the node has. Time a few rows of the study on the surrogate, then set each run’s cores, memory, and wall time from those timings with some margin. A single-threaded model asks for one core per run, and the scheduler packs many runs onto each node.
A job array can also cap how many of its subjobs run at once, which keeps one study from filling a shared queue:
# At most 32 subjobs of this array run at the same time
#PBS -J 1-100%32
#PBS -l select=1:ncpus=1:mem=4gb
Resubmit only the runs that stopped
A long batch on a busy cluster will see some runs stop early: a wall-time limit, a node taken down, an input that needs a fix. Because each run keeps its own status file, resubmitting means finding the runs that have not finished and sending only those:
for run in runs/run-*; do
grep -qx done "$run/status" 2>/dev/null && continue
qsub -v RUN="$(basename "$run")" run-one.pbs
done
A per-run job script that reads RUN from its environment suits this better than an array, because the runs to repeat are rarely a neat range of indices.
Record what produced each run
A result is only useful if someone can tell, months later, what produced it. Write the release version, a checksum of the spreadsheet, and the scheduler’s job ID into every run directory when the job starts:
# At the top of the job script, after cd into the run directory
# ($RELEASE comes from the #PBS -v line in the job script above)
{
echo "release=$("$RELEASE/bin/model" --version)"
echo "study=$(sha256sum "$PBS_O_WORKDIR/study.csv" | cut -d' ' -f1)"
echo "job=$PBS_JOBID"
echo "node=$(hostname)"
} > provenance.txt
With that file beside params.json, any run can be repeated from its own directory, and two batches can be compared knowing exactly which release and which study each came from.
Fall back to one large node without a shared filesystem
Jobs spread over many nodes need their inputs and outputs on storage that every node can reach, such as a shared filesystem. When a cluster has none, the batch can run on one large node instead. That is the single-node fallback we built for this case.
Keep the job script the same and pin every subjob to that node in the resource request. The scheduler still queues the runs, and every run directory stays on one disk:
#PBS -l select=1:ncpus=1:host=bignode
Size the batch to the node’s cores and memory, and let the scheduler keep the queue full. Nothing else in the tooling changes, so the same study can move to the shared-filesystem form later.
Ship the post-run analysis inside the release
A batch produces more output than anyone reads by hand, and an air-gapped network has no cloud tools to help. We shipped post-run analysis that works with no network.
A parser library reads the simulation’s event, scheduler, and catalog outputs. A local dashboard shows eight standard views, and notebooks run the same views for deeper work.
Everything ships in the release’s own Python, so it runs the same on the air-gapped network as on a developer’s machine. Because every run directory holds its params.json, one short step joins the inputs to the results and compares a whole batch:
import json
from pathlib import Path
import pandas as pd
rows = []
for run in sorted(Path("runs").glob("run-*")):
params = json.loads((run / "params.json").read_text())
params["status"] = (run / "status").read_text().strip()
rows.append(params)
batch = pd.DataFrame(rows) # one row per run: inputs and status
Carry the kit in and work on site with the people who run it
The whole kit, with runs, data, and scripts, went onto the air-gapped network by disc. We worked on site with the customer’s staff, as needed and for as long as the work needed. Packing, splitting, and checking a kit like this is covered in Delivering software to air-gapped systems and SCIFs on optical discs.
Results
| Metric | Result |
|---|---|
| Where the model runs | As batch jobs on the customer’s HPC cluster |
| Input for a batch | One design-of-experiments spreadsheet |
| Unit of work | One scheduler job per run, each with its own logs |
| Compiling on the cluster | None |
| Surrogate cluster | Same scheduler; containers, worker count set at launch, workers register themselves |
| Cluster without a shared filesystem | The batch runs on one large node |
| Post-run analysis | Offline dashboard with eight standard views, and notebooks, inside the release |
How we measured: these results come from the delivered tooling and from batches run on the customer’s cluster.
Recommendations
- Build a surrogate with the same scheduler before the first transfer, because every fix there saves a trip through the air gap.
- Package the model so a job needs nothing but the release folder, and test it inside a job on the surrogate.
- Keep the spreadsheet as the only definition of a study, and generate every input file and job script from it.
- Give each run its own directory, logs, and status file, so one run that stops early never hides in the output of another.
- Plan for a cluster without a shared filesystem, and keep the single-node form of the batch ready.
- Ship the analysis with the release, because nothing can be downloaded after the results arrive.
References
- OpenPBS and its source and manual pages
- Docker Compose: docker compose up and –scale
- Python: csv module and string.Template
If your model has outgrown one workstation, see Scientific and HPC software or get a free estimate.