Message boards :
Multicore CPUs :
New batch of QC tasks (QMML)
Message board moderation
Previous · 1 · 2 · 3 · 4 · 5 · 6 . . . 7 · Next
| Author | Message |
|---|---|
|
Send message Joined: 9 Dec 08 Posts: 1006 Credit: 5,068,599 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() |
Can you please confirm that those WUs were cancelled while running and not just while waiting to start? |
|
Send message Joined: 15 Dec 17 Posts: 2 Credit: 5,577,735 RAC: 0 Level ![]() Scientific publications
|
All I can tell is that the one task I had running was a couple hours from completion the last I looked. Then I checked the task list and saw the cancellation: 16776352 12930407 457243 18 Dec 2017 | 6:16:48 UTC 18 Dec 2017 | 14:17:56 UTC Cancelled by server 28,813.53 238,856.00 --- Quantum Chemistry v3.14 (mt) |
|
Send message Joined: 28 Jul 12 Posts: 819 Credit: 1,591,285,971 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
Can you please confirm that those WUs were cancelled while running and not just while waiting to start? On my i7-4770 machine, there were 13 aborted at 13:43:51 UTC. Twelve of them show 0 elapsed time, but the other one shows 05:02:06 (19:52:01) elapsed time. They are all listed as "cancelled by server". And on an i7-3770 machine, three of them completed just after that, at 13:45:09 UTC, after running for around 24 hours or more each, and all show "cancelled by server". Finally, on my Ryzen 1700 machine, two of them completed at 13:52:55 UTC and show "cancelled by server" after running about 18 to 19 hours. So it works. EDIT: But BoincTasks shows the i7-4770 and the Ryzen 1700 machines as "Reported: OK+", so it is only on the GPUGrid status page that the true story is told apparently. |
|
Send message Joined: 9 Dec 08 Posts: 1006 Credit: 5,068,599 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() |
It's true that running tasks are being killed. This is not what I expected. By the way: these WUs should not run 10+ hours on modern CPUs. That's strange. |
|
Send message Joined: 28 Jul 12 Posts: 819 Credit: 1,591,285,971 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
By the way: these WUs should not run 10+ hours on modern CPUs. That's strange. The i7-3770 machine and the Ryzen 1700 machine were running only 2 cores per work unit, while the i7-4770 was running 4 cores per work unit. |
|
Send message Joined: 5 Dec 12 Posts: 84 Credit: 1,663,883,415 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
It's true that running tasks are being killed. This is not what I expected. So far 76 tasks on my machine have been canceled. Please continue killing any task in progress if you don't want the data. No point squandering precious CPU cycles when the science/programming has moved on to a newer revision. Happy Holidays! |
ConanSend message Joined: 25 Mar 09 Posts: 25 Credit: 582,385 RAC: 0 Level ![]() Scientific publications
|
Can you please confirm that those WUs were cancelled while running and not just while waiting to start? Yes I had one that had been running for 33,977 seconds (CPU time 140,002 seconds) and it was cancelled, as well as 2 that had not started. Just an aside to my Fedora 16 Host problems running these work units that are all failing, it is running 'gcc' 4.6. I did some reading on Psi4 and found that it seems to need gcc 4.9 or later in order to run. I have since installed this 'gcc' version on that computer and am awaiting a work unit to see if it works or not. There may still be something missing. I may just have to update Fedora 16 to something more recent. Conan |
|
Send message Joined: 22 Feb 09 Posts: 3 Credit: 114,900 RAC: 0 Level ![]() Scientific publications
|
I have still problems with task miniconda-installer reached time limit 360. Tried 4 tasks today with same result (other 2 task I cancelled). Have standard Fedora 26, nothing special. I don't think the problem is in firewall or slow connection (as suggested in another thread). Is miniconda-installer really downloading something? I rather think, that there is something wrong with installation of files already downloaded on hard-drive. I think I will now wait some time, and come back later (maybe month or two). Hopefully, it will be resolved. |
|
Send message Joined: 10 Sep 10 Posts: 164 Credit: 388,132 RAC: 0 Level ![]() Scientific publications
|
I downloaded 3 wus 3.14 on my vbox linux. They don't start....."Waiting to run". No message on boinc manager. |
ConanSend message Joined: 25 Mar 09 Posts: 25 Credit: 582,385 RAC: 0 Level ![]() Scientific publications
|
I have still problems with task miniconda-installer reached time limit 360. Tried 4 tasks today with same result (other 2 task I cancelled). Have standard Fedora 26, nothing special. Check that SELINUX is not blocking any files from running. I had this problem on my Fedora 25 install and had to create an exception for it. Also make sure your 'gcc' packages are up to date dnf install gcc, or dnf install gcc-c++, should help if you haven't already done so. Conan |
|
Send message Joined: 9 Dec 08 Posts: 1006 Credit: 5,068,599 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() |
@petr - miniconda is downloaded from our servers (~50 MB). After that, at the beginning, psi4 and other packages are downloaded from Anaconda's servers (only the first time). If you suspect a mixup, feel free to "reset" the GPUGRID project and everything should be deleted (and downloaded again at the next WU). Beware that it would kill running tasks! @conan - in principle part of the run is indeed to download its dependencies from Anaconda, including a GCC 5 version which is installed in the project's directory. However, to complete its installation, a library is needed which is generally shipped with... the system's GCC. It's indeed a bit confusing. Maybe your tweak solves the problem, maybe not. |
ConanSend message Joined: 25 Mar 09 Posts: 25 Credit: 582,385 RAC: 0 Level ![]() Scientific publications
|
@ Toni, when these work units run do they get to a certain point then just idle along on a single core for hours on end? I finally got a work unit to download to my Fedora 25 host, and it ran fine up to about 6 hours or so run time and 78.698% completed. After this it has been running on a single core for almost 8 hours now and the progress is still locked at 78.698%. The time to completion has increased from 1 hour 39 minutes to 3 hours 44 minutes and counting. What happened to the Multi-Threading I thought these work units were supposed to do? Run time is now approaching 14 hours and the % done has not moved, this is on a 16 core computer. Conan |
|
Send message Joined: 28 Jul 12 Posts: 819 Credit: 1,591,285,971 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
Run time is now approaching 14 hours and the % done has not moved, this is on a 16 core computer. I seem to recall problems on LHC/ATLAS when running on more than 8 cores, though I was not involved with the problem myself as I run only 7 cores there anyway. But you could try an app_config.xml to limit it to 8 cores. |
|
Send message Joined: 9 Dec 08 Posts: 1006 Credit: 5,068,599 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() |
The computation is done looping over several molecules (~60 if i remember correctly). A checkpoint is written after each loop. Inside a loop there is a part which is multithreaded, and a part which is not. The relative sizes are different. So it's not strange that thread occupancy oscillates. Limiting the number of cores to, say, 4, via the client is ok. |
|
Send message Joined: 2 Jul 16 Posts: 339 Credit: 8,281,341,558 RAC: 2,803 Level ![]() Scientific publications ![]() ![]() ![]() ![]()
|
Run time is now approaching 14 hours and the % done has not moved, this is on a 16 core computer. Someone did tests there. Atlas runs best around 3/4/5 threads. More threads are not utilized very well. I don't recall any mt BOINC app utilizing all threads at 8+. It would probably have to be a straight math project that calculates more #s in parallel to do that. |
|
Send message Joined: 22 Feb 09 Posts: 3 Credit: 114,900 RAC: 0 Level ![]() Scientific publications
|
@Toni, Conan: Thanks to both of you. Looks like the SELINUX was blocking it. It's kind of black box for me. I have found this procedure, which I applied on my system. I hope, that I didn't open the Pandora's box instead :). But after this, the wu started to download additional pkgs and now the computation is running. I will see, if it will succeed, but so far all 6 cores (12 threads) runs at full speed (with some small slowdowns from time to time). So again thx for help. |
ConanSend message Joined: 25 Mar 09 Posts: 25 Credit: 582,385 RAC: 0 Level ![]() Scientific publications
|
@ Toni, when these work units run do they get to a certain point then just idle along on a single core for hours on end? Well I got sick of waiting for this one to finish (it had now been running for over 22 hours still on 1 core) so I created an "app_config.xml" file and inserted that in the project folder and restarted the BOINC Client and Manager. The WU reset itself to 1.098% completed, 5 hours run time and 18 days 23 hours to completion. So 17 to 18 hours run time disappeared and all processing went as well, so apparently no checkpoints. It is now running on 8 cpus instead of 16 which had stopped other work for a day. Will now see what happens. EDIT:: Just after I posted the WU has jumped to 81.741% done and to completion has now dropped to 1 day 11 minutes. So appears to be working heaps better. Conan |
Retvari ZoltanSend message Joined: 20 Jan 09 Posts: 2380 Credit: 16,897,957,044 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
It is now running on 8 cpus instead of 16 which had stopped other work for a day. Your host has 2 CPUs, both have 4 cores hyperthreaded, so the performance scaling will drop rapidly if you run more than 8 threads of Floating Point calculations (most of the science projects are using FP). To all multi-threaded CPU crunchers: Hyperthreaded CPUs have half as many cores as BOINC reports, so you should limit the threads utilized by the app to obtain optimal performance / reliability. |
|
Send message Joined: 5 Dec 12 Posts: 84 Credit: 1,663,883,415 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
If these workunits are gonna take an average of Fifty five hours of CPU time, they really shouldn't crash when I reboot the system. https://www.gpugrid.net/workunit.php?wuid=12932734 Needed to apply updates. Waited until one finished uploading before risking it. Glad I waited. This one was active for less than five minutes. Select Language▼ Follow us on: | Server status | Dayle Diamond [log out] logo About Science Volunteers Performance Forum Join us Donate Name c448-TONI_QMML314rst-0-1-RND2050_0 Workunit 12932734 Created 18 Dec 2017 | 17:32:38 UTC Sent 18 Dec 2017 | 21:17:59 UTC Received 20 Dec 2017 | 1:26:41 UTC Server state Over Outcome Computation error Client state Compute error Exit status 195 (0xc3) EXIT_CHILD_FAILED Computer ID 453935 Report deadline 23 Dec 2017 | 21:17:59 UTC Run time 10.36 CPU time 6.37 Validate state Invalid Credit 0.00 Application version Quantum Chemistry v3.14 (mt) Stderr output <core_client_version>7.8.3</core_client_version> <![CDATA[ <message> process exited with code 195 (0xc3, -61)</message> <stderr_txt> 17:21:56 (70558): wrapper (7.7.26016): starting 17:21:56 (70558): wrapper (7.7.26016): starting 17:21:56 (70558): wrapper: running ../../projects/www.gpugrid.net/Miniconda3-4.3.30-Linux-x86_64.sh (-b -f -p /var/lib/boinc-client/projects/www.gpugrid.net/miniconda) Python 3.6.3 :: Anaconda, Inc. 17:22:04 (70558): miniconda-installer exited; CPU time 6.370986 17:22:04 (70558): wrapper: running /var/lib/boinc-client/projects/www.gpugrid.net/miniconda/bin/python (pre_script.py) CondaValueError: prefix already exists: /var/lib/boinc-client/projects/www.gpugrid.net/miniconda/envs/qmml forrtl: error (78): process killed (SIGTERM) Image PC Routine Line Source libpcm.so.1 00007F3B35B1D725 Unknown Unknown Unknown libpcm.so.1 00007F3B35B1B347 Unknown Unknown Unknown libpcm.so.1 00007F3B35A32AA2 Unknown Unknown Unknown libpcm.so.1 00007F3B35A328F6 Unknown Unknown Unknown libpcm.so.1 00007F3B35A00EFD Unknown Unknown Unknown libpcm.so.1 00007F3B35A04298 Unknown Unknown Unknown libpthread.so.0 00007F3B48D8E150 Unknown Unknown Unknown libmkl_def.so 00007F3B2293D916 Unknown Unknown Unknown forrtl: severe (174): SIGSEGV, segmentation fault occurred Image PC Routine Line Source libpcm.so.1 00007F3B35A04B9A Unknown Unknown Unknown libpthread.so.0 00007F3B48D8E150 Unknown Unknown Unknown Stack trace terminated abnormally. SIGSEGV: segmentation violation Stack trace (11 frames): ../../projects/www.gpugrid.net/wrapper_26198_x86_64-pc-linux-gnu[0x457672] /lib/x86_64-linux-gnu/libpthread.so.0(+0x13150)[0x7fa252998150] ../../projects/www.gpugrid.net/wrapper_26198_x86_64-pc-linux-gnu[0x494313] ../../projects/www.gpugrid.net/wrapper_26198_x86_64-pc-linux-gnu[0x4905b5] /lib/x86_64-linux-gnu/libc.so.6(+0x37140)[0x7fa2525dc140] /lib/x86_64-linux-gnu/libc.so.6(nanosleep+0x58)[0x7fa25267db98] /lib/x86_64-linux-gnu/libc.so.6(usleep+0x44)[0x7fa2526b0134] ../../projects/www.gpugrid.net/wrapper_26198_x86_64-pc-linux-gnu[0x467f2f] ../../projects/www.gpugrid.net/wrapper_26198_x86_64-pc-linux-gnu[0x40b1a1] /lib/x86_64-linux-gnu/libc.so.6(__libc_start_main+0xf1)[0x7fa2525c61c1] ../../projects/www.gpugrid.net/wrapper_26198_x86_64-pc-linux-gnu[0x407ca2] Exiting... 17:24:57 (1324): wrapper (7.7.26016): starting 17:24:57 (1324): wrapper (7.7.26016): starting 17:24:57 (1324): wrapper: running /var/lib/boinc-client/projects/www.gpugrid.net/miniconda/bin/python (pre_script.py) CondaValueError: prefix already exists: /var/lib/boinc-client/projects/www.gpugrid.net/miniconda/envs/qmml CondaHTTPError: HTTP 000 CONNECTION FAILED for url <https://repo.continuum.io/pkgs/main/linux-64/repodata.json.bz2> Elapsed: - An HTTP error occurred when trying to retrieve this URL. HTTP errors are often intermittent, and a simple retry will get you on your way. ConnectionError(MaxRetryError("HTTPSConnectionPool(host='repo.continuum.io', port=443): Max retries exceeded with url: /pkgs/main/linux-64/repodata.json.bz2 (Caused by NewConnectionError('<urllib3.connection.VerifiedHTTPSConnection object at 0x7f19ff5e8898>: Failed to establish a new connection: [Errno -2] Name or service not known',))",),) Traceback (most recent call last): File "pre_script.py", line 13, in <module> raise Exception("Error installing h5py") Exception: Error installing h5py 17:24:58 (1324): $PROJECT_DIR/miniconda/bin/python exited; CPU time 0.257581 17:24:58 (1324): app exit status: 0x1 17:24:58 (1324): called boinc_finish(195) </stderr_txt> ]]> About Science Volunteers Performance Forum Join us Contact Google+ Facebook Twitter © 2017 Universitat Pompeu Fabra |
|
Send message Joined: 13 Dec 17 Posts: 1424 Credit: 9,189,946,190 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]()
|
I finally finished a task so I can post now. Can someone explain what the QC app shows for Status in the BOINC Manager. I had a app_config.xml loaded to limit the number of cpu cores it was supposed to use to 4. However in the Status column it showed 16C for the number of cores allotted. Is that normal? Is that just how it describes itself to BOINC or was it really using all 16 cores? This was my app_confg.xml <app_config> <app> <name>acemdlong</name> <max_concurrent>1</max_concurrent> <gpu_versions> <gpu_usage>1</gpu_usage> <cpu_usage>1</cpu_usage> </gpu_versions> </app> <app> <name>acemdshort</name> <max_concurrent>1</max_concurrent> <gpu_versions> <gpu_usage>1.0</gpu_usage> <cpu_usage>1</cpu_usage> </gpu_versions> </app> <app> <name>QC</name> <max_concurrent>1</max_concurrent> </app> <app_version> <app_name>QC</app_name> <plan_class>mt</plan_class> <avg_ncpus>4</avg_ncpus> <cmdline>--nthreads 4</cmdline> </app_version> </app_config> Does anyone see anything wrong with the app_config? |
©2026 Universitat Pompeu Fabra