New batch of QC tasks (QMML)

Message boards : Multicore CPUs : New batch of QC tasks (QMML)
Message board moderation

To post messages, you must log in.

Previous · 1 · 2 · 3 · 4 · 5 · 6 . . . 7 · Next

AuthorMessage
Toni
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist

Send message
Joined: 9 Dec 08
Posts: 1006
Credit: 5,068,599
RAC: 0
Level
Ser
Scientific publications
watwatwatwat
Message 48415 - Posted: 18 Dec 2017, 18:32:13 UTC - in response to Message 48414.  
Last modified: 18 Dec 2017, 18:32:27 UTC

Can you please confirm that those WUs were cancelled while running and not just while waiting to start?
ID: 48415 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
langfod

Send message
Joined: 15 Dec 17
Posts: 2
Credit: 5,577,735
RAC: 0
Level
Ser
Scientific publications
wat
Message 48416 - Posted: 18 Dec 2017, 18:41:43 UTC - in response to Message 48415.  

All I can tell is that the one task I had running was a couple hours from completion the last I looked.

Then I checked the task list and saw the cancellation:

16776352 12930407 457243
18 Dec 2017 | 6:16:48 UTC 18 Dec 2017 | 14:17:56 UTC
Cancelled by server 28,813.53 238,856.00 ---
Quantum Chemistry v3.14 (mt)
ID: 48416 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Jim1348

Send message
Joined: 28 Jul 12
Posts: 819
Credit: 1,591,285,971
RAC: 0
Level
His
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 48417 - Posted: 18 Dec 2017, 19:11:56 UTC - in response to Message 48415.  
Last modified: 18 Dec 2017, 19:17:36 UTC

Can you please confirm that those WUs were cancelled while running and not just while waiting to start?

On my i7-4770 machine, there were 13 aborted at 13:43:51 UTC. Twelve of them show 0 elapsed time, but the other one shows 05:02:06 (19:52:01) elapsed time. They are all listed as "cancelled by server".

And on an i7-3770 machine, three of them completed just after that, at 13:45:09 UTC, after running for around 24 hours or more each, and all show "cancelled by server".

Finally, on my Ryzen 1700 machine, two of them completed at 13:52:55 UTC and show "cancelled by server" after running about 18 to 19 hours.

So it works.

EDIT: But BoincTasks shows the i7-4770 and the Ryzen 1700 machines as "Reported: OK+", so it is only on the GPUGrid status page that the true story is told apparently.
ID: 48417 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Toni
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist

Send message
Joined: 9 Dec 08
Posts: 1006
Credit: 5,068,599
RAC: 0
Level
Ser
Scientific publications
watwatwatwat
Message 48418 - Posted: 18 Dec 2017, 19:30:51 UTC - in response to Message 48417.  

It's true that running tasks are being killed. This is not what I expected.

By the way: these WUs should not run 10+ hours on modern CPUs. That's strange.
ID: 48418 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Jim1348

Send message
Joined: 28 Jul 12
Posts: 819
Credit: 1,591,285,971
RAC: 0
Level
His
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 48419 - Posted: 18 Dec 2017, 20:59:20 UTC - in response to Message 48418.  

By the way: these WUs should not run 10+ hours on modern CPUs. That's strange.

The i7-3770 machine and the Ryzen 1700 machine were running only 2 cores per work unit, while the i7-4770 was running 4 cores per work unit.
ID: 48419 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Dayle Diamond

Send message
Joined: 5 Dec 12
Posts: 84
Credit: 1,663,883,415
RAC: 0
Level
His
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 48420 - Posted: 18 Dec 2017, 22:21:25 UTC

It's true that running tasks are being killed. This is not what I expected.


So far 76 tasks on my machine have been canceled.
Please continue killing any task in progress if you don't want the data.
No point squandering precious CPU cycles when the science/programming has moved on to a newer revision.

Happy Holidays!
ID: 48420 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Profile Conan

Send message
Joined: 25 Mar 09
Posts: 25
Credit: 582,385
RAC: 0
Level
Gly
Scientific publications
wat
Message 48421 - Posted: 19 Dec 2017, 3:17:38 UTC - in response to Message 48415.  

Can you please confirm that those WUs were cancelled while running and not just while waiting to start?


Yes I had one that had been running for 33,977 seconds (CPU time 140,002 seconds) and it was cancelled, as well as 2 that had not started.

Just an aside to my Fedora 16 Host problems running these work units that are all failing, it is running 'gcc' 4.6.
I did some reading on Psi4 and found that it seems to need gcc 4.9 or later in order to run.

I have since installed this 'gcc' version on that computer and am awaiting a work unit to see if it works or not. There may still be something missing.

I may just have to update Fedora 16 to something more recent.

Conan
ID: 48421 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Petr Kriz

Send message
Joined: 22 Feb 09
Posts: 3
Credit: 114,900
RAC: 0
Level

Scientific publications
wat
Message 48422 - Posted: 19 Dec 2017, 8:05:25 UTC

I have still problems with task miniconda-installer reached time limit 360. Tried 4 tasks today with same result (other 2 task I cancelled). Have standard Fedora 26, nothing special.
I don't think the problem is in firewall or slow connection (as suggested in another thread). Is miniconda-installer really downloading something? I rather think, that there is something wrong with installation of files already downloaded on hard-drive.
I think I will now wait some time, and come back later (maybe month or two). Hopefully, it will be resolved.
ID: 48422 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
[VENETO] boboviz

Send message
Joined: 10 Sep 10
Posts: 164
Credit: 388,132
RAC: 0
Level

Scientific publications
wat
Message 48424 - Posted: 19 Dec 2017, 9:16:06 UTC

I downloaded 3 wus 3.14 on my vbox linux.
They don't start....."Waiting to run". No message on boinc manager.
ID: 48424 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Profile Conan

Send message
Joined: 25 Mar 09
Posts: 25
Credit: 582,385
RAC: 0
Level
Gly
Scientific publications
wat
Message 48425 - Posted: 19 Dec 2017, 10:28:30 UTC - in response to Message 48422.  

I have still problems with task miniconda-installer reached time limit 360. Tried 4 tasks today with same result (other 2 task I cancelled). Have standard Fedora 26, nothing special.
I don't think the problem is in firewall or slow connection (as suggested in another thread). Is miniconda-installer really downloading something? I rather think, that there is something wrong with installation of files already downloaded on hard-drive.
I think I will now wait some time, and come back later (maybe month or two). Hopefully, it will be resolved.


Check that SELINUX is not blocking any files from running. I had this problem on my Fedora 25 install and had to create an exception for it.

Also make sure your 'gcc' packages are up to date
dnf install gcc, or dnf install gcc-c++, should help if you haven't already done so.

Conan
ID: 48425 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Toni
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist

Send message
Joined: 9 Dec 08
Posts: 1006
Credit: 5,068,599
RAC: 0
Level
Ser
Scientific publications
watwatwatwat
Message 48426 - Posted: 19 Dec 2017, 10:32:19 UTC - in response to Message 48424.  
Last modified: 19 Dec 2017, 10:33:32 UTC

@petr - miniconda is downloaded from our servers (~50 MB). After that, at the beginning, psi4 and other packages are downloaded from Anaconda's servers (only the first time). If you suspect a mixup, feel free to "reset" the GPUGRID project and everything should be deleted (and downloaded again at the next WU). Beware that it would kill running tasks!

@conan - in principle part of the run is indeed to download its dependencies from Anaconda, including a GCC 5 version which is installed in the project's directory. However, to complete its installation, a library is needed which is generally shipped with... the system's GCC. It's indeed a bit confusing. Maybe your tweak solves the problem, maybe not.
ID: 48426 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Profile Conan

Send message
Joined: 25 Mar 09
Posts: 25
Credit: 582,385
RAC: 0
Level
Gly
Scientific publications
wat
Message 48428 - Posted: 19 Dec 2017, 12:16:24 UTC
Last modified: 19 Dec 2017, 12:22:40 UTC

@ Toni, when these work units run do they get to a certain point then just idle along on a single core for hours on end?

I finally got a work unit to download to my Fedora 25 host, and it ran fine up to about 6 hours or so run time and 78.698% completed.
After this it has been running on a single core for almost 8 hours now and the progress is still locked at 78.698%.
The time to completion has increased from 1 hour 39 minutes to 3 hours 44 minutes and counting.

What happened to the Multi-Threading I thought these work units were supposed to do?

Run time is now approaching 14 hours and the % done has not moved, this is on a 16 core computer.

Conan
ID: 48428 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Jim1348

Send message
Joined: 28 Jul 12
Posts: 819
Credit: 1,591,285,971
RAC: 0
Level
His
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 48429 - Posted: 19 Dec 2017, 12:47:20 UTC - in response to Message 48428.  

Run time is now approaching 14 hours and the % done has not moved, this is on a 16 core computer.

I seem to recall problems on LHC/ATLAS when running on more than 8 cores, though I was not involved with the problem myself as I run only 7 cores there anyway. But you could try an app_config.xml to limit it to 8 cores.
ID: 48429 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Toni
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist

Send message
Joined: 9 Dec 08
Posts: 1006
Credit: 5,068,599
RAC: 0
Level
Ser
Scientific publications
watwatwatwat
Message 48430 - Posted: 19 Dec 2017, 13:14:27 UTC - in response to Message 48428.  

The computation is done looping over several molecules (~60 if i remember correctly). A checkpoint is written after each loop. Inside a loop there is a part which is multithreaded, and a part which is not. The relative sizes are different. So it's not strange that thread occupancy oscillates. Limiting the number of cores to, say, 4, via the client is ok.
ID: 48430 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
mmonnin

Send message
Joined: 2 Jul 16
Posts: 339
Credit: 8,281,341,558
RAC: 2,803
Level
Tyr
Scientific publications
watwatwatwatwat
Message 48431 - Posted: 19 Dec 2017, 14:25:37 UTC - in response to Message 48429.  

Run time is now approaching 14 hours and the % done has not moved, this is on a 16 core computer.

I seem to recall problems on LHC/ATLAS when running on more than 8 cores, though I was not involved with the problem myself as I run only 7 cores there anyway. But you could try an app_config.xml to limit it to 8 cores.


Someone did tests there. Atlas runs best around 3/4/5 threads. More threads are not utilized very well. I don't recall any mt BOINC app utilizing all threads at 8+. It would probably have to be a straight math project that calculates more #s in parallel to do that.
ID: 48431 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Petr Kriz

Send message
Joined: 22 Feb 09
Posts: 3
Credit: 114,900
RAC: 0
Level

Scientific publications
wat
Message 48433 - Posted: 19 Dec 2017, 16:20:43 UTC

@Toni, Conan: Thanks to both of you. Looks like the SELINUX was blocking it. It's kind of black box for me. I have found this procedure, which I applied on my system. I hope, that I didn't open the Pandora's box instead :).
But after this, the wu started to download additional pkgs and now the computation is running. I will see, if it will succeed, but so far all 6 cores (12 threads) runs at full speed (with some small slowdowns from time to time).
So again thx for help.
ID: 48433 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Profile Conan

Send message
Joined: 25 Mar 09
Posts: 25
Credit: 582,385
RAC: 0
Level
Gly
Scientific publications
wat
Message 48438 - Posted: 19 Dec 2017, 22:18:59 UTC - in response to Message 48428.  
Last modified: 19 Dec 2017, 22:36:38 UTC

@ Toni, when these work units run do they get to a certain point then just idle along on a single core for hours on end?

I finally got a work unit to download to my Fedora 25 host, and it ran fine up to about 6 hours or so run time and 78.698% completed.
After this it has been running on a single core for almost 8 hours now and the progress is still locked at 78.698%.
The time to completion has increased from 1 hour 39 minutes to 3 hours 44 minutes and counting.

What happened to the Multi-Threading I thought these work units were supposed to do?

Run time is now approaching 14 hours and the % done has not moved, this is on a 16 core computer.

Conan


Well I got sick of waiting for this one to finish (it had now been running for over 22 hours still on 1 core) so I created an "app_config.xml" file and inserted that in the project folder and restarted the BOINC Client and Manager.

The WU reset itself to 1.098% completed, 5 hours run time and 18 days 23 hours to completion.
So 17 to 18 hours run time disappeared and all processing went as well, so apparently no checkpoints.

It is now running on 8 cpus instead of 16 which had stopped other work for a day.
Will now see what happens.

EDIT:: Just after I posted the WU has jumped to 81.741% done and to completion has now dropped to 1 day 11 minutes.
So appears to be working heaps better.


Conan
ID: 48438 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Profile Retvari Zoltan
Avatar

Send message
Joined: 20 Jan 09
Posts: 2380
Credit: 16,897,957,044
RAC: 0
Level
Trp
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 48440 - Posted: 20 Dec 2017, 1:28:12 UTC - in response to Message 48438.  

It is now running on 8 cpus instead of 16 which had stopped other work for a day.
Will now see what happens.

Your host has 2 CPUs, both have 4 cores hyperthreaded, so the performance scaling will drop rapidly if you run more than 8 threads of Floating Point calculations (most of the science projects are using FP).
To all multi-threaded CPU crunchers: Hyperthreaded CPUs have half as many cores as BOINC reports, so you should limit the threads utilized by the app to obtain optimal performance / reliability.
ID: 48440 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Dayle Diamond

Send message
Joined: 5 Dec 12
Posts: 84
Credit: 1,663,883,415
RAC: 0
Level
His
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 48441 - Posted: 20 Dec 2017, 1:32:50 UTC

If these workunits are gonna take an average of Fifty five hours of CPU time, they really shouldn't crash when I reboot the system.

https://www.gpugrid.net/workunit.php?wuid=12932734

Needed to apply updates. Waited until one finished uploading before risking it. Glad I waited. This one was active for less than five minutes.


Select Language​▼

Twitter
Facebook
Follow us on:
|
Server status
|
Dayle Diamond [log out]

logo
About Science Volunteers Performance Forum Join us Donate

Name c448-TONI_QMML314rst-0-1-RND2050_0
Workunit 12932734
Created 18 Dec 2017 | 17:32:38 UTC
Sent 18 Dec 2017 | 21:17:59 UTC
Received 20 Dec 2017 | 1:26:41 UTC
Server state Over
Outcome Computation error
Client state Compute error
Exit status 195 (0xc3) EXIT_CHILD_FAILED
Computer ID 453935
Report deadline 23 Dec 2017 | 21:17:59 UTC
Run time 10.36
CPU time 6.37
Validate state Invalid
Credit 0.00
Application version Quantum Chemistry v3.14 (mt)
Stderr output

<core_client_version>7.8.3</core_client_version>
<![CDATA[
<message>
process exited with code 195 (0xc3, -61)</message>
<stderr_txt>
17:21:56 (70558): wrapper (7.7.26016): starting
17:21:56 (70558): wrapper (7.7.26016): starting
17:21:56 (70558): wrapper: running ../../projects/www.gpugrid.net/Miniconda3-4.3.30-Linux-x86_64.sh (-b -f -p /var/lib/boinc-client/projects/www.gpugrid.net/miniconda)
Python 3.6.3 :: Anaconda, Inc.
17:22:04 (70558): miniconda-installer exited; CPU time 6.370986
17:22:04 (70558): wrapper: running /var/lib/boinc-client/projects/www.gpugrid.net/miniconda/bin/python (pre_script.py)

CondaValueError: prefix already exists: /var/lib/boinc-client/projects/www.gpugrid.net/miniconda/envs/qmml

forrtl: error (78): process killed (SIGTERM)
Image PC Routine Line Source
libpcm.so.1 00007F3B35B1D725 Unknown Unknown Unknown
libpcm.so.1 00007F3B35B1B347 Unknown Unknown Unknown
libpcm.so.1 00007F3B35A32AA2 Unknown Unknown Unknown
libpcm.so.1 00007F3B35A328F6 Unknown Unknown Unknown
libpcm.so.1 00007F3B35A00EFD Unknown Unknown Unknown
libpcm.so.1 00007F3B35A04298 Unknown Unknown Unknown
libpthread.so.0 00007F3B48D8E150 Unknown Unknown Unknown
libmkl_def.so 00007F3B2293D916 Unknown Unknown Unknown
forrtl: severe (174): SIGSEGV, segmentation fault occurred
Image PC Routine Line Source
libpcm.so.1 00007F3B35A04B9A Unknown Unknown Unknown
libpthread.so.0 00007F3B48D8E150 Unknown Unknown Unknown

Stack trace terminated abnormally.
SIGSEGV: segmentation violation
Stack trace (11 frames):
../../projects/www.gpugrid.net/wrapper_26198_x86_64-pc-linux-gnu[0x457672]
/lib/x86_64-linux-gnu/libpthread.so.0(+0x13150)[0x7fa252998150]
../../projects/www.gpugrid.net/wrapper_26198_x86_64-pc-linux-gnu[0x494313]
../../projects/www.gpugrid.net/wrapper_26198_x86_64-pc-linux-gnu[0x4905b5]
/lib/x86_64-linux-gnu/libc.so.6(+0x37140)[0x7fa2525dc140]
/lib/x86_64-linux-gnu/libc.so.6(nanosleep+0x58)[0x7fa25267db98]
/lib/x86_64-linux-gnu/libc.so.6(usleep+0x44)[0x7fa2526b0134]
../../projects/www.gpugrid.net/wrapper_26198_x86_64-pc-linux-gnu[0x467f2f]
../../projects/www.gpugrid.net/wrapper_26198_x86_64-pc-linux-gnu[0x40b1a1]
/lib/x86_64-linux-gnu/libc.so.6(__libc_start_main+0xf1)[0x7fa2525c61c1]
../../projects/www.gpugrid.net/wrapper_26198_x86_64-pc-linux-gnu[0x407ca2]

Exiting...
17:24:57 (1324): wrapper (7.7.26016): starting
17:24:57 (1324): wrapper (7.7.26016): starting
17:24:57 (1324): wrapper: running /var/lib/boinc-client/projects/www.gpugrid.net/miniconda/bin/python (pre_script.py)

CondaValueError: prefix already exists: /var/lib/boinc-client/projects/www.gpugrid.net/miniconda/envs/qmml


CondaHTTPError: HTTP 000 CONNECTION FAILED for url <https://repo.continuum.io/pkgs/main/linux-64/repodata.json.bz2>
Elapsed: -

An HTTP error occurred when trying to retrieve this URL.
HTTP errors are often intermittent, and a simple retry will get you on your way.
ConnectionError(MaxRetryError("HTTPSConnectionPool(host='repo.continuum.io', port=443): Max retries exceeded with url: /pkgs/main/linux-64/repodata.json.bz2 (Caused by NewConnectionError('<urllib3.connection.VerifiedHTTPSConnection object at 0x7f19ff5e8898>: Failed to establish a new connection: [Errno -2] Name or service not known',))",),)


Traceback (most recent call last):
File "pre_script.py", line 13, in <module>
raise Exception("Error installing h5py")
Exception: Error installing h5py
17:24:58 (1324): $PROJECT_DIR/miniconda/bin/python exited; CPU time 0.257581
17:24:58 (1324): app exit status: 0x1
17:24:58 (1324): called boinc_finish(195)

</stderr_txt>
]]>

About
Science
Volunteers
Performance
Forum
Join us
Contact

Google+ Facebook Twitter
© 2017 Universitat Pompeu Fabra
ID: 48441 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Keith Myers
Avatar

Send message
Joined: 13 Dec 17
Posts: 1424
Credit: 9,189,946,190
RAC: 0
Level
Tyr
Scientific publications
watwatwatwatwat
Message 48444 - Posted: 20 Dec 2017, 4:23:12 UTC

I finally finished a task so I can post now.

Can someone explain what the QC app shows for Status in the BOINC Manager. I had a app_config.xml loaded to limit the number of cpu cores it was supposed to use to 4. However in the Status column it showed 16C for the number of cores allotted.

Is that normal? Is that just how it describes itself to BOINC or was it really using all 16 cores?

This was my app_confg.xml
<app_config>
<app>
<name>acemdlong</name>
<max_concurrent>1</max_concurrent>
<gpu_versions>
<gpu_usage>1</gpu_usage>
<cpu_usage>1</cpu_usage>
</gpu_versions>
</app>
<app>
<name>acemdshort</name>
<max_concurrent>1</max_concurrent>
<gpu_versions>
<gpu_usage>1.0</gpu_usage>
<cpu_usage>1</cpu_usage>
</gpu_versions>
</app>
<app>
<name>QC</name>
<max_concurrent>1</max_concurrent>
</app>
<app_version>
<app_name>QC</app_name>
<plan_class>mt</plan_class>
<avg_ncpus>4</avg_ncpus>
<cmdline>--nthreads 4</cmdline>
</app_version>
</app_config>


Does anyone see anything wrong with the app_config?
ID: 48444 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Previous · 1 · 2 · 3 · 4 · 5 · 6 . . . 7 · Next

Message boards : Multicore CPUs : New batch of QC tasks (QMML)

©2026 Universitat Pompeu Fabra