Message boards :
Multicore CPUs :
Simultaneously starting MCs
Message board moderation
Previous · 1 · 2 · 3 · Next
| Author | Message |
|---|---|
|
Send message Joined: 17 Feb 09 Posts: 91 Credit: 1,603,303,394 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
Started two QC Wu's simultaneously using app version 3.20 and one failed with this stderr report: <core_client_version>7.9.3</core_client_version> <![CDATA[ <message> process exited with code 195 (0xc3, -61)</message> <stderr_txt> 13:45:21 (5569): wrapper (7.7.26016): starting 13:45:21 (5569): wrapper (7.7.26016): starting 13:45:21 (5569): wrapper: running /bin/bash (-c "flock /var/lib/boinc/projects/www.gpugrid.net/miniconda.lock ./miniconda-installer.sh -b -u -p /var/lib/boinc/projects/www.gpugrid.net/miniconda") flock: failed to execute ./miniconda-installer.sh: Text file busy 13:45:30 (5569): /bin/bash exited; CPU time 0.000265 13:45:30 (5569): app exit status: 0x45 13:45:30 (5569): called boinc_finish(195) </stderr_txt> ]]> Machine is an AMD FX8350 at 4.1 GHz with Fedora 28 kernel 4.16.13 with both gcc and glibc-devel installed. I had been running both my 8 core machines in 4 core mode and 2 concurrent QC WU's for several months with only one simultaneous start error the whole time but with no other project to compete for the CPU. I found that when running both WCG and QC CPU WU's, simultaneous starts occurred more frequently and unfortunately when it did happen, then failures would leap frog through the queue quickly and blacklist the computer from more work for a day. |
|
Send message Joined: 9 Dec 08 Posts: 1006 Credit: 5,068,599 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() |
Yes, now simultaneous starts crash with "text file busy". Will look for yet another workaround. |
|
Send message Joined: 9 Dec 08 Posts: 1006 Credit: 5,068,599 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() |
Attempting fix at version 321 |
|
Send message Joined: 21 Mar 16 Posts: 513 Credit: 4,673,458,277 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
Just got a couple of errors from last night's WUs version 3.20. Here are the links: https://www.gpugrid.net/result.php?resultid=17740487 https://www.gpugrid.net/result.php?resultid=17740592 All the rest ran fine. |
|
Send message Joined: 9 Dec 08 Posts: 1006 Credit: 5,068,599 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() |
^^ These seem connection errors (network down or so) |
|
Send message Joined: 21 Mar 16 Posts: 513 Credit: 4,673,458,277 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
^^ These seem connection errors (network down or so) Network can affect computation after the WU is already downloaded? |
|
Send message Joined: 3 Sep 14 Posts: 152 Credit: 927,557,369 RAC: 132 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
I have exactly the same problems so I dont think this is connection related. |
|
Send message Joined: 9 May 13 Posts: 171 Credit: 4,739,796,466 RAC: 1,441 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
Version 321 looks promising. Just finished two that started at the same time and they finished normally. Just started three at the same time and they are all processing as they should. |
|
Send message Joined: 17 Feb 09 Posts: 91 Credit: 1,603,303,394 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
Success with two QC WU's simultaneous start using app 3.21. Both WU's happily crunching @ 1.098% completed so far. Event Log Excerpt: Thu 07 Jun 2018 07:22:04 AM MST | GPUGRID | task m0000000638_2278bdac_n00020-SDOERR_QMML50-0-1-RND5957_0 resumed by user Thu 07 Jun 2018 07:22:04 AM MST | GPUGRID | task m0000000643_5ed133e9_n00020-SDOERR_QMML50-0-1-RND3172_0 resumed by user Thu 07 Jun 2018 07:22:05 AM MST | GPUGRID | Starting task m0000000638_2278bdac_n00020-SDOERR_QMML50-0-1-RND5957_0 Thu 07 Jun 2018 07:22:05 AM MST | GPUGRID | Starting task m0000000643_5ed133e9_n00020-SDOERR_QMML50-0-1-RND3172_0 |
|
Send message Joined: 9 Dec 08 Posts: 1006 Credit: 5,068,599 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() |
^^ These seem connection errors (network down or so) Yes. WUs check the latest version of conda packages/libraries right after start (from conda cloud). |
|
Send message Joined: 2 Jul 16 Posts: 339 Credit: 8,281,341,558 RAC: 3,417 Level ![]() Scientific publications ![]() ![]() ![]() ![]()
|
^^ These seem connection errors (network down or so) Dang is that only after start? If tasks were started, paused, another started, etc could networking be disabled after they start? Or at each start/resume? Not me, but I've heard of some setup schedules to allow downloads at certain times of the day due to varying bandwidth costs. |
|
Send message Joined: 17 Feb 09 Posts: 91 Credit: 1,603,303,394 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
Today I have had three successful simultaneous starts and returns without error on two different FX8350 machines. As far as I am concerned, this bug has been resolved at least for the Fedora distro. |
|
Send message Joined: 21 Mar 16 Posts: 513 Credit: 4,673,458,277 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
Haven't gotten a single error with version 3.21 and it's been running all day. What did you change? |
|
Send message Joined: 9 Dec 08 Posts: 1006 Credit: 5,068,599 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() |
The main change was locking the miniconda directory upon initial installation/update. This in turn required some workarounds. May not be perfect but should be much better. Regarding network access, it is attempted at each WU start or re-start. The amount of downloaded data should be usually negligible (except the first time). Edit to add: the network accesses are only to "conda cloud", a python distribution and package manager. |
|
Send message Joined: 10 Sep 10 Posts: 164 Credit: 388,132 RAC: 0 Level ![]() Scientific publications
|
No more error, but a strange behaviour. Remanining time for every wus are over 20h, but wus are crunched in 40/50 minutes... |
|
Send message Joined: 13 Dec 17 Posts: 1424 Credit: 9,189,946,190 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]()
|
Didn't get any response from my post in the cpu tasks thread. How do you get the cpu tasks? I tried and failed. Is the QC app still considered a Test app? That was the only Preference toggle I didn't select. |
|
Send message Joined: 23 Dec 09 Posts: 189 Credit: 4,813,881,008 RAC: 181 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
Did you check: use CPU as well? You might not have allowed it. No it is not necessary to check BETA tasks. I do not have checked it either, but I do get QC tasks. |
|
Send message Joined: 13 Dec 17 Posts: 1424 Credit: 9,189,946,190 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]()
|
Yes CPU was checked for use both places and the QC app. Didn't get any QC tasks both time I tried. Just gpu tasks. The scheduler request was for both cpu and gpu work. I know there is plenty of cpu tasks to farm out. Couldn't explain why I didn't get any work. I'll have to try again without gpu work checked I guess. |
|
Send message Joined: 28 Jul 12 Posts: 819 Credit: 1,591,285,971 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
Yes CPU was checked for use both places and the QC app. Didn't get any QC tasks both time I tried. Just gpu tasks. That is your problem. The BOINC scheduler can get all mixed up when you select both CPU and GPU work on the same project, and you are eventually left high and dry on one or the other. There are several discussions on it at Einstein, which has the same problem since they do both CPU and GPU work. Here is one recent discussion, where the moderator explains why the requester is not getting GPU work. https://einsteinathome.org/content/not-getting-gpu-wus-anymore#comment-165295 I use separate machines for the CPU work and the GPU work on GPUGrid. |
|
Send message Joined: 13 Dec 17 Posts: 1424 Credit: 9,189,946,190 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]()
|
Thanks for the reply. I don't know. It worked a couple of months ago when the QC app and tasks first showed up. I was crunching both gpu and cpu at the same time. I know that shutting off a gpu request will probably work just to get some of the new QC tasks along with the latest 3.21 app. I was just wondering how well the app works now with concurrent starts. |
©2026 Universitat Pompeu Fabra