Hi there. Let me ask you here about a little help...
I downloaded "ThornList-ET_2018_02" from:
https://bitbucket.org/zach_etienne/wvuthorns_diagnostics/src/master/
then I run "./GetComponents ThornList-ET_2018_02" and I've got an output as you can see in "output.dat". It created a Cactus directory for me, so I went there and tried to compile it by:
./simfactory/bin/sim setup-silent
./simfactory/bin/sim build -j12 --thornlist ../ThornList-ET_2018_02 --optionlist ../dns4.cfg
My dns4.cfg option list is probably irrelevant, but I'm sending it too. The compilation failed, as you can see in x4x.out.
Any idea? Any simple and straightforward solution?
PS: Let me also mention that when I try to get components again, I have an info related to these particular components which are mentioned in the compilation errors. Look at the output2.dat.
Michal Pirog
Hi all,
Here are the minutes from today’s meeting.
Chair: Peter Minutes: Leo Present: Peter Diener, Leo Werneck, Sam Cupp,
Zach Etienne, Keith Dow, Steve Brandt, Beyhan, Gabriele Bozzola,
Roland Haas, Yosef Zlochower
* Updates on new thorns proposed for inclusion and revisions:
- NRPyEllipticET: got nice feedback from reviewers. Leo will be
working on updates over the next couple of weeks.
- SelfForce1D: no progress during the past week.
- Baikal: Updates are mostly going to be applied to the code
generation scripts.
- FLRW Solver: thorn has been thoroughly reviewed for best
programming practices and is looking good.
- Canuda Solver: No update since no reviewer or champion present.
* Zach announces: CarpetX-compatible versions of Baikal and IllinoisGRMHD
will be developed. "BaikalX" may already exist, will look into it.
* Steve asks: can NRPyElliptic handle the surfaces of NSs when generating
ID? Zach says: BNS initial data will be the subject of a future project.
There is no doubt that NRPyElliptic can solve the elliptic problem, but
the techniques requires for handling the NS surface are yet not
completely clear.
* Zach suggests: some of ADMBase's functionalities (e.g., setting ID) should
be moved elsewhere, as currently the thorn does (too) much more than what
its name suggests.
Cheers,
Leo
------
Leonardo R. Werneck, Ph.D.
Postdoctoral researcher
Office EP 314 | Department of Physics | University of Idaho
875 Perimeter Dr. MS 0903
Moscow, ID 83844-0903, USA
leonardo(a)uidaho.edu <mailto:leonardo@uidaho.edu>
https://leowerneck.github.io <https://leowerneck.github.io/>
Hello,
Please consider joining the weekly Einstein Toolkit phone call at
9:00 am US central time on Thursdays. For details on how to connect
and what agenda items are to be discussed, use the link below.
https://docs.einsteintoolkit.org/et-docs/Main_Page#Weekly_Users_Call
--The Maintainers
Hi all,
I am trying to perform simulations on the Discoverer cluster (
https://docs.discoverer.bg/index.html) using the latest Einstein Toolkit
(ETK) release (ET_2022_05) and the Spritz GRMHD code.
To compile ETK on Discoverer, I am attaching the simfactory configuration
files which I had newly prepared. I am also attaching the list of modules
which were loaded.
To submit the simulation, for instance, I use the following simfactory
command:
sim submit BNS_IF_fluxCT_dx018_q10_RPA_RotGas_E8e49
--parfile=./par/BNS_IF_fluxCT_dx018_q10_RPA_RotGas_E8e49.par
--config=newspritzgnu --machine=discoverer --procs=256 --num-threads=1
--ppn-used=128 --walltime=24:00:00
Unfortunately, my simulations crash after running for some time. My guess
is that I might not be correctly setting the configuration options or flags
during compilation, which might affect my simulation during runtime, but I
am not completely certain.
I am attaching the output file, the error file, the parfile as well as the
generated backtrace for the simulation which used 256 procs. I also looked
at the hexadecimal addresses in the backtrace with addr2line, but
unfortunately all of them return "??:0"
I also noticed that when changing the number of processors, the simulation
crashes at different times. But if I keep the number of processors as the
same, the simulation always crashes at the same point.
For instance, simulation with 256 processors ran for about 2 hours on the
cluster, and crashed after completing about 4600 iterations. One submitted
with 1280 processors ran for about 12 hours and crashed after completing
about 32000 iterations. Simulation with 1792 processors instead crashed
soon after the start of the simulation (within few minutes), even before
reaching iteration 0. For all cases, I always set number of threads as 1.
If you have any suggestions or insights on why the simulations crash and in
case I have any incorrect settings in the configuration files, kindly let
me know. I would greatly appreciate your help. If you need any further
information from my side, please let me know too.
Thank you very much.
Kind regards,
Jay Kalinani
Present: Roland, Peter, Keith, Zach, Beyhan, Leo, Gabriele, Yosef
new ET modules:
* Zach reviewed FLRW solver, thumbs up for does no harm
** Roland to look at file not found error during compile
* no news on Canuda review
* Peter is reviewing SelfForce1D, thumbs up for harm
* NRPyEllipticET, Leo provided reviewers with parfile
** Leo to contact reviewers
Functionality to retire:
* keep ongoing list in wiki
* Roland will push pull request that remove use of "REQUIRES THORNS"
from current ET thorns
Questions on mailing list:
* nothing new
Open tickets:
* Zach reported
https://bitbucket.org/einsteintoolkit/tickets/issues/2635 that
contains a compile time warning about working only with Cartesian
grid, suggests removal
* New ticket https://bitbucket.org/einsteintoolkit/tickets/issues/2634/
mentioning lack of documentation on CCTK_BUILTIN_EXPECT
* Two tickets on SummationByParts assigned to Peter,
https://bitbucket.org/einsteintoolkit/tickets/issues/2633 and
https://bitbucket.org/einsteintoolkit/tickets/issues/2632
* Roland filed a ticket about non-conforming handling of initial data
selection parameters
https://bitbucket.org/einsteintoolkit/tickets/issues/2631 . Zach
provided some recap that setting shift to 0 rather than using
LORENE's shift gives better evolution, in particular much less
eccentricity. Gabriele suggests adding a comment about this to the
gallery example BNS parfile, instead of using LORENE's shift.
* Zach and Gabriele suggest adding warning about this to code, also
suggest adding similar runtime warning for other settings that we
expect to produce failures or strange results
* Gabriele brought
up https://bitbucket.org/einsteintoolkit/tickets/issues/2629 . Had a
discussion on how to handle scheduling to compute a quantity every N
steps. Roland to send email about what he recollects about doing this
in EVOL vs ANALYSIS to the mailing list.
No ET call next week due to overlap with EU ET summer school.
chair next time: Peter
minutes next time: Leo
Yours,
Roland
--
My email is as private as my paper mail. I therefore support encrypting
and signing email messages. Get my PGP key from http://pgp.mit.edu .
Hello,
Please consider joining the weekly Einstein Toolkit phone call at
9:00 am US central time on Thursdays. For details on how to connect
and what agenda items are to be discussed, use the link below.
** DAYLIGHT SAVING TIME WARNING **
Please note that the US / EU has already / not yet transitioned to /
from daylight saving time. The phone call will be at 15:30 Central EU
time.
https://docs.einsteintoolkit.org/et-docs/Main_Page#Weekly_Users_Call
--The Maintainers
Hi Roland and Peter,
No worries, thank you for the follow-up! I added the memory directive and received the same segfault error (and the backtrace looks the same as well). Each node on Quartz has 512GB for 128 cores, so I requested 4GB per core. For example, a submission with 10 cores (2 processes with 5 threads each) had:
#SBATCH --mem=40GB
I repeated the attempt with a different setup but the same memory scaling of 4 GB per core, such as 2 processes with 1 thread each (so 8GB memory requested), and failed the same way. Also tried using an entire node (8 processes with 16 threads per) and all its memory (--mem=0), and failed.
I'm using the default static_tov.par, no changes:
IO::out_dir = $parfile
IOScalar::outScalar_every = 32
IOScalar::one_file_per_group = yes
IOScalar::outScalar_vars = "
HydroBase::rho
HydroBase::press
HydroBase::eps
HydroBase::vel
ADMBase::lapse
ADMBase::metric
ADMBase::curv
ML_ADMConstraints::ML_Ham
ML_ADMConstraints::ML_mom
"
Thank you,
Jessica
Dr. Jessica S. Warren
Physics Lecturer
Indiana University Northwest
warrenjs(a)iun.edu
________________________________
From: Roland Haas
Sent: Thursday, August 18, 2022 3:36 PM
To: Warren, Jessica Sawyer
Cc: users(a)einsteintoolkit.org
Subject: Re: [Users] [External] Re: Running with SLURM
Hello Jessica,
Sorry for the delay in responding.
We discussed your issue during today's weekly Einstein Toolkit call
(http://lists.einsteintoolkit.org/pipermail/users/2022-August/008660.html)
and there were some suggestions.
The error you see is somewhat puzzling. The segfault happens in a MPI
call that is part of the Carpet Reduction call (during output).
The puzzling thing is that this is not the first MPI call in this run
(there are some much earlier) nor the first reduction (eg there are
reductions for the min/max output that you saw on screen).
There was one suggestions that were mentioned:
* Peter Diener noted that he had seem on one cluster issues with errors
and segfaults later in the run where he needed to explicitly pass a
"#SBATCH --memory XGB" to sbatch to request that memory is available
>From the fact that you can see output to screen but the failure in a
reduction to me sounds like the issue is somehow encountered while
executing code in the CarpetIOScalar thorn (IOBasic is to screen,
IOScalar is to disk). Are you passing any "strange" options or special
variables to its outScalar_vars option?
Yours,
Roland
> Hi Roland,
>
> The admins reinstalled openmpi and it now runs the hello script
> correctly. However, the Toolkit would still produce seg faults after
> srun. Switching to mvapich seems to have largely done the trick
> though, as the TOV job is now able to start executing. As long as
> there is only 1 MPI process (with however many threads), the TOV job
> runs to completion correctly. However, anytime there are multiple
> MPI processes, it crashes at the first time iteration:
>
> INFO (TOVSolver): Done interpolation.
> ---------------------------------------------------------------------------
> Iteration Time | ADMBASE::alp |
> HYDROBASE::rho | minimum maximum | minimum maximum
> ---------------------------------------------------------------------------
> 0 0.000 | 0.6698612 0.9966374 | 1.000000e-10
> 0.0012800 Rank 1 with PID 3964893 received signal 11
> Writing backtrace to static_tov/backtrace.1.txt
> srun: error: c40: task 1: Segmentation fault (core dumped)
>
> The backtrace is attached, as well as the last portion of the output,
> and it looks like the issue is tied to Carpet. Are there some
> settings in the parameter file that need adjusting or setting to fix
> this? Or perhaps specific settings for the number of ranks and
> threads?
>
> Thank you,
> Jessica
>
>
> Dr. Jessica S. Warren
> Physics Lecturer
> Indiana University Northwest
> warrenjs(a)iun.edu
>
> ________________________________
> From: Roland Haas
> Sent: Thursday, August 11, 2022 8:32 AM
> To: Warren, Jessica Sawyer
> Cc: users(a)einsteintoolkit.org
> Subject: Re: [Users] [External] Re: Running with SLURM
>
> Hello Jessica,
>
> If you get the same error from hello-world and from Cactus then it
> would seem that there is still something off with the MPI stack.
>
> The -lmpi_cxx option instructs the linker to link in C++ bindings for
> MPI though for just the hello world example, it being C code, this is
> not required and -lmpi alone is sufficient.
>
> I would see two options that would let you get running somewhat
> quickly:
>
> 1. report your issues with OpenMPI and hello-world (including link to
> the source code on the web, and the exact command line to compile) to
> the admins and ask them for help
>
> 1.5 instead of using gcc to compile for OpenMPI do use the MPI
> official compiler wrapper mpicc which would just be:
>
> mpicc -o hello hello.c
>
> that is you do not have to pass and library or inlcude options. If
> this fails, I would definitely talk to the admins.
>
> 2. compile hello-world using mvapich. For this the easiest way is to
> make sure to load the mvapich module and then use the same compiler
> wrapper invication to compile:
>
> mpicc -o hello hello.c
>
> If 2 works then you can also compile the Einstein Toolkit with
> mvapich. You have to make sure to load the correct module before
> compiling the toolkit and then ExternalLibraries/MPI should figure
> out (from the mpicc wrapper) how to compile the toolkit.
>
> Yours,
> Roland
>
>
> > Hi Roland,
> >
> > Thank you so much. The compute nodes are able to be used for
> > compilation, and the directories match what is listed in
> > make.MPI.defn. When doing the 'hello' example you linked to, it was
> > unable to compile due to a linker error (/usr/bin/ld: cannot find
> > -lmpi_cxx). I re-ran it in verbose mode and found the directory it
> > was searching did exist and did have lmpi but not lmpi_cxx. The
> > admins said they had had some issues installing openmpi (couldn't
> > recall exactly what), and recommended mpavich (since that does have
> > lmpicxx installed and is their preferred implementation). However,
> > they reinstalled openmpi in an effort to get that to work and it did
> > allow the 'hello' script to compile, but when executed it produced:
> >
> > --------------------------------------------------------------------------
> > No OpenFabrics connection schemes reported that they were able to be
> > used on a specific port. As such, the openib BTL (OpenFabrics
> > support) will be disabled for this port.
> >
> > Local host: h1
> > Local device: mlx5_0
> > Local port: 1
> > CPCs attempted: rdmacm, udcm
> > --------------------------------------------------------------------------
> > Hello world from processor h1.quartz.uits.iu.edu, rank 0 out of 1
> > processors
> >
> > Similarly, doing the TOV job via sbatch, after the srun command it
> > gave the same OpenFabrics message (for each MPI rank) and then the
> > same segmentation faults as before. I've contacted the admins about
> > this and am waiting to hear back. Do you have any recommendations -
> > perhaps it would be easier to try switching over to mvapich? If so,
> > could you point me to some resources on how to reconfigure?
> >
> > Thank you,
> > Jessica
> >
> > Dr. Jessica S. Warren
> > Physics Lecturer
> > Indiana University Northwest
> > warrenjs(a)iun.edu
> > ________________________________
> > From: Roland Haas <rhaas(a)illinois.edu>
> > Sent: Tuesday, August 9, 2022 9:48 AM
> > To: Warren, Jessica Sawyer <warrenjs(a)iun.edu>
> > Cc: users(a)einsteintoolkit.org <users(a)einsteintoolkit.org>
> > Subject: [External] Re: [Users] Running with SLURM
> >
> > Hello Jessica,
> >
> > You may also find something useful in the setting up a new machine
> > seminar presentation:
> >
> > https://urldefense.com/v3/__https://www.einsteintoolkit.org/seminars/2022_0…
> >
> > Yours,
> > Roland
> >
> > --
> > My email is as private as my paper mail. I therefore support
> > encrypting and signing email messages. Get my PGP key from
> > https://urldefense.com/v3/__http://pgp.mit.edu__;!!DZ3fjg!9JAgxc4juluJwklwT…
> > .
>
>
> --
> My email is as private as my paper mail. I therefore support
> encrypting and signing email messages. Get my PGP key from
> https://urldefense.com/v3/__http://pgp.mit.edu__;!!DZ3fjg!_ZQHbCvNiX5H7WOd1…
> .
--
My email is as private as my paper mail. I therefore support encrypting
and signing email messages. Get my PGP key from http://pgp.mit.edu .
Hello all,
As discussed and voted on during the ET call yesterday, the weekly ET
call time has been changed (back) to:
Thu 9:00am US Central time
Reminders, website and wiki have been updated accordingly.
Yours,
Roland
--
My email is as private as my paper mail. I therefore support encrypting
and signing email messages. Get my PGP key from http://pgp.mit.edu .