Hi Ian and Erik,
Thank you very much for all the advice and pointers so far!
I didn't compile the ET myself; it was done by an HPC engineer. He is unfamiliar with Cactus and started off not using a config file, so he had to troubleshoot his way through the compilation process. We are both scratching our heads about what the issue with mpirun could be.
I suspect he didn't set MPI_DIR, so I'm going to suggest that he fixes that and see if recompiling takes care of things.
The scheduler automatically terminates jobs that run on too many processors. For my simulation, this appears to happen as soon as TwoPunctures starts generating the initial data. I then get error messages of the form: "Job terminated as it used more cores (17.6) than requested (4)." (I switched from requesting 3 processors to requesting 4.) The number of cores it tries to use appears to differ from run to run.
The parameter file uses Carpet. It generates the following output (when I request 4 processors):
INFO (Carpet): MPI is enabled
INFO (Carpet): Carpet is running on 4 processes
INFO (Carpet): This is process 0
INFO (Carpet): OpenMP is enabled
INFO (Carpet): This process contains 16 threads, this is thread 0
INFO (Carpet): There are 64 threads in total
INFO (Carpet): There are 16 threads per process
Mpirun gives me the following information for the node allocation: slots=4, max_slots=0, slots_inuse=0, state=UP.
The tree view of the processes looks like this:
PID TTY STAT TIME COMMAND
19503 ? S 0:00 sshd: allgwy001@pts/7
19504 pts/7 Ss 0:00 \_ -bash
6047 pts/7 R+ 0:00 \_ ps -u allgwy001 f
Adding "cat $PBS_NODEFILE" to my PBS script didn't seem to produce anything, although I could be doing something stupid. I'm very new to the syntax!
Erik: documentation for the cluster I'm trying to run on (called HEX) is available under the Documentation tab at
hpc.uct.ac.za, but it's very basic.
Thanks again for all your help! I'll let you know if we make any progress.
Gwyneth