Message boards :
Multicore CPUs :
New batch of QC tasks (QMML)
Message board moderation
Previous · 1 . . . 3 · 4 · 5 · 6 · 7 · Next
| Author | Message |
|---|---|
|
Send message Joined: 21 Mar 16 Posts: 513 Credit: 4,673,458,277 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
50,000 WUs... Holy ****. If only those were GPU WUs |
|
Send message Joined: 2 Jul 16 Posts: 339 Credit: 8,281,341,558 RAC: 2,803 Level ![]() Scientific publications ![]() ![]() ![]() ![]()
|
50,000 WUs... Holy ****. If only those were GPU WUs I saw that too and downloaded 4 on a computer I haven't run any of these on. Started at the same time and boom all went to crap. I even tried to pause some to stop the error. http://www.gpugrid.net/results.php?hostid=458003 |
|
Send message Joined: 17 Feb 09 Posts: 91 Credit: 1,603,303,394 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
So far from what I have learned starting more than one multiple cpu job at a time in a split core senerio is that they need not be started at exactly the same time for one to error. Given not an exact start at the same time, the one started first always errors and the second one started processes to successful completion with the next WU in the queue started to also complete successfully. Four of four tries, near simultaneous starts the one started first ended up failing. Secondly, controlling the time to when the second WU is allowed to start following the first start time is up to 5 seconds as tested so far. Third, when the boinc client switches between projects, the QC WU's so far observed are completed in pairs leaving no single job left in progress (suspended) to stagger start the times. This unfortunate characteristic means "simultaneous (or nearly so) starts" cause an error whenever the client switches back to the gpugrid cpu jobs with a queue larger than one WU. Guess the only way to prevent this behavior is to not split cores, especially for unattended clients. Edit: correct my lousy spelling, stupid keyboard :) |
|
Send message Joined: 13 Dec 17 Posts: 1424 Credit: 9,189,946,190 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]()
|
I didn't have that experience with the two TONI tasks I started simultaneously. Or within the 5 second window you described. Both completed successfully. I am limiting core usage to four with an app_config file. Limiting the max_concurrent to 1 now since I also crunch SETI cpu tasks on that computer. I ran the two concurrent jobs when Toni requested users to try that experiment. |
|
Send message Joined: 13 Dec 17 Posts: 1424 Credit: 9,189,946,190 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]()
|
Wow, the credit awarded is all over the place for these DOMINIK tasks. Obviously NOT tied to computation time or resources used for compute. Task 16862848 3715 seconds CPU time Credit awarded 21 Task 16815159 3687 seconds CPU time Credit awarded 161 |
|
Send message Joined: 29 Dec 16 Posts: 2 Credit: 1,397 RAC: 0 Level ![]() Scientific publications
|
I just got a task and it finished on a dual e5-2450l 32g ram server fine. resultid=16815237 Keep developing this, please. This is quite nice and I would really like to see this as part of GPUGRID permanently. |
|
Send message Joined: 29 Dec 16 Posts: 2 Credit: 1,397 RAC: 0 Level ![]() Scientific publications
|
resultid=16815470 |
|
Send message Joined: 15 Dec 17 Posts: 9 Credit: 0 RAC: 0 Level ![]() Scientific publications
|
Hello Keith, Wow, the credit awarded is all over the place for these DOMINIK tasks. Obviously NOT tied to computation time or resources used for compute. Did you observe this behavior multiple times? Really strange to be honest. Thanks for helping out everyone! |
|
Send message Joined: 2 Jul 16 Posts: 339 Credit: 8,281,341,558 RAC: 2,803 Level ![]() Scientific publications ![]() ![]() ![]() ![]()
|
All tasks were erroring out in 2 minutes due to the app using gcc5.5 This got it to go to farther and start being multithreaded. We'll see if it actually completes. sudo apt-get install gcc-5 g++-5 |
|
Send message Joined: 2 Jul 16 Posts: 339 Credit: 8,281,341,558 RAC: 2,803 Level ![]() Scientific publications ![]() ![]() ![]() ![]()
|
Yup, it completed. http://www.gpugrid.net/result.php?resultid=16817115 |
|
Send message Joined: 15 Dec 17 Posts: 9 Credit: 0 RAC: 0 Level ![]() Scientific publications
|
Great! Thank you very much |
|
Send message Joined: 23 Dec 09 Posts: 189 Credit: 4,813,881,008 RAC: 149 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
@klepel - can you try installing gcc (if not already there)? I tried it yesterday. I installed gcc-5 and gcc-6. And it worked on the computer http://www.gpugrid.net/results.php?hostid=452211 |
|
Send message Joined: 17 Feb 09 Posts: 91 Credit: 1,603,303,394 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
These are the two WU's that were started about 5 seconds apart on an AMD FX-8350 with the first one started failing. Stdoutdea.txt: 03-Jan-2018 17:20:14 [GPUGRID] [css] running e113s22_e86s4p0f123-PABLO_p53_PHEX10P_IDP-0-1-RND2720_0 (0.987 CPUs + 1 NVIDIA GPU) 03-Jan-2018 17:20:14 [GPUGRID] Starting task c00000_00024-DOMINIK_QMML2_m0000000055-0-1-RND3244_0 03-Jan-2018 17:20:14 [GPUGRID] [cpu_sched] Starting task c00000_00024-DOMINIK_QMML2_m0000000055-0-1-RND3244_0 using QC version 314 (mt) in slot 9 03-Jan-2018 17:20:14 [GPUGRID] [css] running c00000_00024-DOMINIK_QMML2_m0000000055-0-1-RND3244_0 (4 CPUs) 03-Jan-2018 17:20:20 [GPUGRID] task c06475_06499-DOMINIK_QMML2_m0000000054-0-1-RND3067_0 resumed by user 03-Jan-2018 17:20:21 [GPUGRID] [css] running e113s22_e86s4p0f123-PABLO_p53_PHEX10P_IDP-0-1-RND2720_0 (0.987 CPUs + 1 NVIDIA GPU) 03-Jan-2018 17:20:21 [GPUGRID] [css] running c00000_00024-DOMINIK_QMML2_m0000000055-0-1-RND3244_0 (4 CPUs) 03-Jan-2018 17:20:21 [GPUGRID] Starting task c06475_06499-DOMINIK_QMML2_m0000000054-0-1-RND3067_0 03-Jan-2018 17:20:21 [GPUGRID] [cpu_sched] Starting task c06475_06499-DOMINIK_QMML2_m0000000054-0-1-RND3067_0 using QC version 314 (mt) in slot 10 03-Jan-2018 17:20:21 [GPUGRID] [css] running c06475_06499-DOMINIK_QMML2_m0000000054-0-1-RND3067_0 (4 CPUs) 03-Jan-2018 17:20:25 [GPUGRID] [sched_op] Deferring communication for 00:01:39 03-Jan-2018 17:20:25 [GPUGRID] [sched_op] Reason: Unrecoverable error for task c00000_00024-DOMINIK_QMML2_m0000000055-0-1-RND3244_0 03-Jan-2018 17:20:25 [GPUGRID] Computation for task c00000_00024-DOMINIK_QMML2_m0000000055-0-1-RND3244_0 finished 03-Jan-2018 17:20:25 [GPUGRID] [css] running e113s22_e86s4p0f123-PABLO_p53_PHEX10P_IDP-0-1-RND2720_0 (0.987 CPUs + 1 NVIDIA GPU) 03-Jan-2018 17:20:25 [GPUGRID] [css] running c06475_06499-DOMINIK_QMML2_m0000000054-0-1-RND3067_0 (4 CPUs) I didn't have that experience with the two TONI tasks I started simultaneously. Or within the 5 second window you described. Both completed successfully. I am limiting core usage to four with an app_config file. Limiting the max_concurrent to 1 now since I also crunch SETI cpu tasks on that computer. I ran the two concurrent jobs when Toni requested users to try that experiment. Since this hasn't been the case with your Intel's implies this could be a cpu related phenomenom (architecture/scheduling differences). Perhaps the Intel's can handle initial start up processes faster than the FX series AMD, (might spring for a Ryen7 soon just to check them as well). Regardless, the issue is resolved with my systems by limiting concurrent QC jobs to one and use the other four cores to run WCG as to date I have not experienced a concurrent issue with the WCG WU's. |
|
Send message Joined: 13 Dec 17 Posts: 1424 Credit: 9,189,946,190 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]()
|
Hello Keith, Yes, the first completed tasks got reasonable credit. Then when I downloaded more, all the credit for them nosedived. Once I saw that they weren't worth running I set NNT. 16862848 12962574 456812 3 Jan 2018 | 23:38:10 UTC 4 Jan 2018 | 1:20:46 UTC Completed and validated 1,000.69 3,715.60 21.12 Quantum Chemistry v3.14 (mt) 16815333 12963093 456812 4 Jan 2018 | 1:00:00 UTC 4 Jan 2018 | 2:57:45 UTC Completed and validated 1,020.00 3,812.63 27.38 Quantum Chemistry v3.14 (mt) 16815332 12963092 456812 4 Jan 2018 | 0:59:23 UTC 4 Jan 2018 | 2:40:47 UTC Completed and validated 990.38 3,697.44 25.78 Quantum Chemistry v3.14 (mt) 16815320 12963080 456812 4 Jan 2018 | 1:00:37 UTC 4 Jan 2018 | 3:14:28 UTC Completed and validated 997.66 3,731.61 27.28 Quantum Chemistry v3.14 (mt) 16815307 12963067 456812 4 Jan 2018 | 1:07:01 UTC 4 Jan 2018 | 4:04:01 UTC Completed and validated 1,009.28 3,642.48 26.40 Quantum Chemistry v3.14 (mt) 16815275 12963035 456812 4 Jan 2018 | 1:07:38 UTC 4 Jan 2018 | 4:21:09 UTC Completed and validated 1,033.63 3,668.31 26.52 Quantum Chemistry v3.14 (mt) 16815264 12963024 456812 4 Jan 2018 | 1:06:24 UTC 4 Jan 2018 | 3:46:54 UTC Completed and validated 969.15 3,592.69 25.72 Quantum Chemistry v3.14 (mt) 16815248 12963008 456812 4 Jan 2018 | 0:58:46 UTC 4 Jan 2018 | 2:24:17 UTC Completed and validated 935.31 3,503.52 23.59 Quantum Chemistry v3.14 (mt) 16815234 12962994 456812 4 Jan 2018 | 1:05:48 UTC 4 Jan 2018 | 3:30:44 UTC Completed and validated 981.30 3,616.45 26.43 Quantum Chemistry v3.14 (mt) 16815171 12962931 456812 3 Jan 2018 | 23:37:35 UTC 4 Jan 2018 | 1:04:10 UTC Completed and validated 929.82 3,523.44 18.45 Quantum Chemistry v3.14 (mt) 16815159 12962919 456812 3 Jan 2018 | 23:35:43 UTC 4 Jan 2018 | 0:16:22 UTC Completed and validated 986.88 3,687.96 161.36 Quantum Chemistry v3.14 (mt) |
|
Send message Joined: 23 Dec 09 Posts: 189 Credit: 4,813,881,008 RAC: 149 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
I have to report back on the AMD Ryzen 1700x Computer: http://www.gpugrid.net/results.php?hostid=420971 If I run 3 instances, the WUs crashes after about 200 seconds, and after that the computer crashes completely. If I am running one (01) instance (WU), the computer runs without any problem. However, as BOINC downloads several of this Quantum Chemistry v3.14 (mt) WUs, BOINC thinks my CPU cache is full and refuses to download additional CPU WUs from PRIMGRID. So after a while the CPU is only loaded with one QC WU (4 threads) and the rest of the cores are idle - Not very efficient. Sorry, Dominik unter diesen Umständen kann ich keine weiteren QC WUs für diesen Computer herunterladen. Komme aber gerne zurück, wenn wir ohne Probleme mehrere MultiCores WUs gleichzeitig bearbeiten können. |
|
Send message Joined: 17 Feb 09 Posts: 91 Credit: 1,603,303,394 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
However, as BOINC downloads several of this Quantum Chemistry v3.14 (mt) WUs, BOINC thinks my CPU cache is full and refuses to download additional CPU WUs from PRIMGRID. So after a while the CPU is only loaded with one QC WU (4 threads) and the rest of the cores are idle - Not very efficient. If you temporarily suspend the QC jobs (except perhaps the one in progress), you should download more work from your other projects and once downloaded, resume the QC jobs and let boinc take over running the various projects as you have them configured. You may need to "update" the projects you want more work from under the "projects" tab to initiate the downloads right away. Edit: Close quote and add last sentence. |
|
Send message Joined: 13 Dec 17 Posts: 1424 Credit: 9,189,946,190 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]()
|
Seti@home has been down all day so ran out of work. Decided to give the QC Dominik tasks another try. Thought that possibly the last batch were unique or one-offs or something different about the computer from when I first ran them. Nope. Even worse credit for the batch I ran this afternoon. Credits awarded = 6. If you expect to get anybody to want to run these, you are going to have to make them more appealing, credit-wise at least. For me, not worth the electricity to run them. Would rather let the computer go cold and give my power bill a temporary reprieve. |
|
Send message Joined: 2 Jul 16 Posts: 339 Credit: 8,281,341,558 RAC: 2,803 Level ![]() Scientific publications ![]() ![]() ![]() ![]()
|
Yeah credit took a dump on the last two I completed. Run Time-----CPU Time-----Credit 2,389.43-----26,206.90-----549.62 3,034.92-----36,710.33-----94.32 |
|
Send message Joined: 16 May 13 Posts: 41 Credit: 145,731,947 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]()
|
I got 5.97 points for 32 minutes calculation on my fx 6100. That's ridiculous! |
|
Send message Joined: 13 Dec 17 Posts: 1424 Credit: 9,189,946,190 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]()
|
If you care to learn about why the low credit or why certain tasks get hi-middle-low credit, read my post here |
©2026 Universitat Pompeu Fabra