Sorry, fixed the typo (though error is still there - having tried different parallel environments). I've attached a chunk of the output but killed it before it carried on until error. Do you need further output?
The double comment is ignored - will confirm though...
On 20 October 2015 at 13:24, Ian Hinder ian.hinder@aei.mpg.de wrote:
On 20 Oct 2015, at 13:44, Geraint Pratten g.pratten@sussex.ac.uk wrote:
Hi,
So I'm getting a problem passing the correct number of MPI processes through to Carpet. I'm not sure what I've messed up in the configurations, the error I am getting is:
The environment variable CACTUS_NUM_PROCS is set to 4, but there are 1 MPI processes. This may indicate a severe problem with the MPI startup mechanism.
I've attached the .run and .sub scripts that I use. Can anyone see anything obviously wrong? The local cluster uses the UNIVA Grid Engine for submission scripts.
Hi Geraint,
In your submission script, you have
#! /bin/bash #$ -cwd #$ -j y #$ -o output/@SIMULATION_NAME@.out #$ -e output/@SIMULATION_NAME@.err
## Error and ouput #$ -m abe #$ -M geraint.pratten@gmail.com
## Set up parallel environment ##$ -pe openmpi_mixed_32 @PROCS_REQUESTED@ #$ -pe mpich @PROC_REQUESTEDS@
## Request queue and job class #$ -q mps.q #$ -jc mps.medium
## Walltime, simulation name and memory #$ -l h_rt=@WALLTIME@ #$ -N @SHORT_SIMULATION_NAME@
Specifically, you say @PROC_REQUESTED@ instead of @PROCS_REQUESTED@. i.e. a missing S on PROCS. Can you also post your output and error files? It would be good if simfactory could detect errors like this. It would also be good if the queuing system on that machine would complain if it got a number of procs that it doesn't understand. Maybe it does complain, but it's only a warning?
I'm also not sure what happens if you do a double comment ##. It may treat it as a single comment, and use the first pe that you specify.
-- Ian Hinder http://members.aei.mpg.de/ianhin
On 20 Oct 2015, at 14:49, Geraint Pratten g.pratten@sussex.ac.uk wrote:
Sorry, fixed the typo (though error is still there - having tried different parallel environments). I've attached a chunk of the output but killed it before it carried on until error. Do you need further output?
Hi Geraint,
In apollo.run, you call "mpirun". This might not be the right mpirun. It might be better to call it using $MPIDIR/bin/mpirun, if that is where it is located.
From the output file, you can see that there are multiple Cactus banners appearing. This indicates that each mpi process is starting thinking it is the only one there. This is a smoking gun for a mismatch between the MPI versions used to compile and run, so I would focus on that possibility.
users@lists.einsteintoolkit.org