Hi
Please consider joining the weekly Einstein Toolkit phone call at
10 am US central time on Mondays. As usual, you can find instructions
how to join on the following web site:
http://einsteintoolkit.org/community/support/
In short: the number is (+1) 225-578-4942 or (+1) 866-573-0359 and
the conference id is 118682#.
Some tickets within the last week that are still open:
- 1618: adaptive MPI (to be reviewed)
- 1518: Parameter Parser (still in review)
- 1648: MPI config (review)
- 1651: order or expansion in parameter files
Other projects include
- Make Illinois code work with standard ET
- git transition
As always, please add to this list if you like.
Frank Loeffler
Hi everyone,
I am trying to configure and add a new machine (cluster) into my
Simfactory machine database. (.../simfactory/mdb/machinses/).
However, I do not know what exactly what the following variables or
abbreviations are in the .ini file: spn, mpn. I have looked up values
given to these in different machine files and found that they are not
the same for all machines. Which brings me to the question of what
they are so I can give to them the correct values relevant to/for the
new machine that I want to add.
I understand ppn represents the number *p*rocessors *p*er *n*ode or
number of cores per node (SimfactoryAdvancedTutorial)
<https://docs.einsteintoolkit.org/et-docs/Simulation_Factory_Advanced_Tutoriā¦>
. On this
one, what exactly should determine the value for *min-ppn* as far
as the machine being defined is concerned? Or the min-ppn's
value doesn't depend on the machine?
Thank you in advance fr your help.
Dumsani
Hi all,
Has anyone run into problems recently with Cactus jobs on Stampede? I've had jobs die when checkpointing, and also mysteriously hanging for no apparent reason. These might be separate problems. The checkpointing issue occurred when I submitted several jobs and they all started checkpointing at the same time after 3 hours. The hang happened after a few hours of evolution, with GDB reporting
> MPIDI_CH3I_MRAILI_Get_next_vbuf (vc_ptr=0x7fff00d9a8d8, vbuf_ptr=0x13)
> at src/mpid/ch3/channels/mrail/src/gen2/ibv_channel_manager.c:296
> 296 for (; i < mv2_MPIDI_CH3I_RDMA_Process.polling_group_size;
> ++i)
Unfortunately I didn't ask for a backtrace. I'm using mvapich2. I've been in touch with support and they said the dying while checkpointing coincided with the filesystems being hit hard by my jobs, which makes sense, but they didn't see any problems in their logs, and they have no idea about the mysterious hang. I repeated the hanging job and it ran fine.
--
Ian Hinder
http://numrel.aei.mpg.de/people/hinder