Message boards :
Multicore CPUs :
New batch of QC tasks (QMML)
Message board moderation
Previous · 1 · 2 · 3 · 4 · 5 . . . 7 · Next
| Author | Message |
|---|---|
|
Send message Joined: 16 May 13 Posts: 41 Credit: 145,731,947 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]()
|
I'm not getting any WUs, neither on my AMD, nor on my Intel CPUs. |
|
Send message Joined: 9 Dec 08 Posts: 1006 Credit: 5,068,599 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() |
Let me summarize the current status. We are making tests in view of a large production run. The WUs which are out now are called QMML314long and last several hours. This longer test have a couple of new failure modes which I think are related to restarts, and can be fixed. Another different problem is task distribution by the BOINC scheduler. First of all, as said above, some hosts are ignored for no reason I can fathom. Another is that some hosts are "soaking up" dozens of WUs, which means they are not available to others. I am hoping that both problems will sort out by themselves with a sufficiently large batch. Final notes: (a) CPU maker is irrelevant. (b) disappeared WUs were tests which I cancelled from the server. |
ConanSend message Joined: 25 Mar 09 Posts: 25 Credit: 582,385 RAC: 0 Level ![]() Scientific publications
|
The Multiple threaded work units that were sent out last month worked fine for me with no issues. Just an update as I am still am getting these errors, this is the full error CondaValueError: prefix already exists: /home/xxxxxxxx/BOINC/projects/www.gpugrid.net/miniconda/envs/qmml ERROR conda.core.link:_execute_actions(337): An error occurred while installing package 'psi4::gcc-5-5.2.0-1'. LinkError: post-link script failed for package psi4::gcc-5-5.2.0-1 running your command again with `-v` will provide additional information location of failed script: /home/xxxxxxxx/BOINC/projects/www.gpugrid.net/miniconda/envs/qmml/bin/.gcc-5-post-link.sh ==> script messages <== <None> Attempting to roll back. LinkError: post-link script failed for package psi4::gcc-5-5.2.0-1 running your command again with `-v` will provide additional information location of failed script: /home/xxxxxxxx/BOINC/projects/www.gpugrid.net/miniconda/envs/qmml/bin/.gcc-5-post-link.sh ==> script messages <== <None> Traceback (most recent call last): File "pre_script.py", line 20, in <module> raise Exception("Error installing psi4 dev") Exception: Error installing psi4 dev 10:18:33 (23979): $PROJECT_DIR/miniconda/bin/python exited; CPU time 69.668408 10:18:33 (23979): app exit status: 0x1 10:18:33 (23979): called boinc_finish(195) Conan |
|
Send message Joined: 9 Dec 08 Posts: 1006 Credit: 5,068,599 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() |
@conan: do you have "gcc" installed in your system? If not, can you try to install it? |
ConanSend message Joined: 25 Mar 09 Posts: 25 Credit: 582,385 RAC: 0 Level ![]() Scientific publications
|
@conan: do you have "gcc" installed in your system? If not, can you try to install it? It was installed on on one computer with Fedora 25, but was not installed on the other two with Fedora 16 and Fedora 21, all 64 bit. Have installed now and await to see what happens. Versions range from 4.6.3-2 (Fedora 16), 4.9.2-6 (Fedora 21) to 6.4.1-1 (Fedora 25). Thanks Conan |
|
Send message Joined: 28 Jul 12 Posts: 819 Credit: 1,591,285,971 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
Another different problem is task distribution by the BOINC scheduler. First of all, as said above, some hosts are ignored for no reason I can fathom. Another is that some hosts are "soaking up" dozens of WUs, which means they are not available to others. I am hoping that both problems will sort out by themselves with a sufficiently large batch. Yes! I just got some QC on my Ryzen 1700. All good things come to those that wait. (The first four errored out after a couple of minutes, but the fifth one is running fine after 50 minutes and I think it will fly, running two cores on each WU.) |
|
Send message Joined: 9 Dec 08 Posts: 1006 Credit: 5,068,599 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() |
I may have understood the problem of hosts not getting WUs. I was sending tasks at a high priority, which means they crossed the threshold to only go to "reliable hosts" -- a questionable heuristic. 100 tasks named "s*-QMML314long" I made at a lower priority seem to have been sent quickly. |
|
Send message Joined: 16 May 13 Posts: 41 Credit: 145,731,947 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]()
|
I got my first WU today. Unfortunately the WU needs 4,7 GB of ram. Can you optimise that? |
ConanSend message Joined: 25 Mar 09 Posts: 25 Credit: 582,385 RAC: 0 Level ![]() Scientific publications
|
@conan: do you have "gcc" installed in your system? If not, can you try to install it? The Fedora 16 host still has the same error, but the Fedora 21 host is processing a work unit now for the last 8 hours 21 minutes and 68% done, so it looks good at this point. My Fedora 25 host has not received any work yet so can't say about that one. My WU is using 1.5 GB of RAM. Thanks Conan |
|
Send message Joined: 23 Dec 09 Posts: 189 Credit: 4,813,881,008 RAC: 149 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
I may have understood the problem of hosts not getting WUs. I was sending tasks at a high priority, which means they crossed the threshold to only go to "reliable hosts" -- a questionable heuristic. This solved it for my second computer. Works on a USB Stick with Lubuntu 17.04. Unfortunatelly, crashed: http://www.gpugrid.net/result.php?resultid=16776102 |
ConanSend message Joined: 25 Mar 09 Posts: 25 Credit: 582,385 RAC: 0 Level ![]() Scientific publications
|
@conan: do you have "gcc" installed in your system? If not, can you try to install it? This WU on the Fedora 21 host worked and completed successfully, my first of this batch. Conan |
|
Send message Joined: 9 Dec 08 Posts: 1006 Credit: 5,068,599 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() |
@klepel - can you try installing gcc (if not already there)? tks |
|
Send message Joined: 5 Dec 12 Posts: 84 Credit: 1,663,883,415 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
https://drive.google.com/file/d/1bKmSXT4IAVTR8b-fpiGdC6Gm4szduk0X/view?usp=sharing Running one now. 1950x, Linux 17.10. Average time taken is two hours, fifteen minutes per task. At the time of the screenshot, the work unit is around fourty percent done. I'm watching my CPU usage hit 100%, stay there for a while, then...waves. I don't think it's thermal throttling. It's not overclocked and WCG tasks only make those patterns when tasks are starting/finishing. If it's working as intended, ok. |
ConanSend message Joined: 25 Mar 09 Posts: 25 Credit: 582,385 RAC: 0 Level ![]() Scientific publications
|
What are the requirements for the running of Psi4? My Fedora 16 computer after installing "gcc" is still getting the same error that it failed to install, but the Fedora 21 computer is now running fine. Is there a certain "glibc", "gcc" or Linux kernel that is required to install this programme? My older Fedora 16 install may not meet the requirements perhaps? I still can't get any work on my Intel Xeon running Fedora 25, keeps saying that there is no work available when in fact there is, but that is another issue. Thanks Conan |
|
Send message Joined: 9 Dec 08 Posts: 1006 Credit: 5,068,599 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() |
@dayle - oscillating CPU% is expected and due to the parts of the calculation which are not parallelized. Thermal throttling is unlikely imho (and I imagine it would manifests as a decrease in CPU clock, not CPU%). @conan - in principle a system with GLIBC>=2.14 should be capable to run; Fedora 16 seemed to have it but it is old, so probably something else is missing. Sorry. I've made another thread with information which may be useful. |
|
Send message Joined: 16 May 13 Posts: 41 Credit: 145,731,947 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]()
|
My WU is using 1.5 GB of RAM. 18-Dec-2017 14:07:22 [GPUGRID] Quantum Chemistry needs 4768.37 MB RAM but only 3469.24 MB is available for use. |
|
Send message Joined: 9 Dec 08 Posts: 1006 Credit: 5,068,599 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() |
We are talking about 3 different memory use figures: A. The amount of memory actually used (which varies with time), which Conan measuread as 1.5 GB B. The amount "requested" by the workunit, currently 4 GB C. The maximum amount your boinc client allows to use (you can configure this to some extent) The following should hold: A < B < C In your case, B>C and therefore the WU was not allowed to start (I guess). |
|
Send message Joined: 28 Jul 12 Posts: 819 Credit: 1,591,285,971 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
@dayle - oscillating CPU% is expected and due to the parts of the calculation which are not parallelized. Thermal throttling is unlikely imho (and I imagine it would manifests as a decrease in CPU clock, not CPU%). That one is a little confusing. The remaining time also increases as a consequence. I think I aborted some unnecessarily when it appeared that they were stuck. I now just let them run. Maybe you should make a big sticky on it to catch people's attention? |
|
Send message Joined: 15 Dec 17 Posts: 2 Credit: 5,577,735 RAC: 0 Level ![]() Scientific publications
|
Did tasks just get aborted by the system? Name s51-TONI_QMML314long-0-1-RND1523_1 Workunit 12930407 Exit status 202 (0xca) EXIT_ABORTED_BY_PROJECT <core_client_version>7.8.3</core_client_version> <![CDATA[ <message> aborted by project - no longer usable</message> <stderr_txt> |
|
Send message Joined: 28 Jul 12 Posts: 819 Credit: 1,591,285,971 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
Did tasks just get aborted by the system? Yes. I just had a bunch aborted at 13:44 UTC. But there are now new ones in the pipeline. |
©2026 Universitat Pompeu Fabra