New batch of QC tasks (QMML)

Message boards : Multicore CPUs : New batch of QC tasks (QMML)
Message board moderation

To post messages, you must log in.

Previous · 1 . . . 3 · 4 · 5 · 6 · 7 · Next

AuthorMessage
PappaLitto

Send message
Joined: 21 Mar 16
Posts: 513
Credit: 4,673,458,277
RAC: 0
Level
Arg
Scientific publications
watwatwatwatwatwatwatwat
Message 48587 - Posted: 3 Jan 2018, 23:39:37 UTC

50,000 WUs... Holy ****. If only those were GPU WUs
ID: 48587 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
mmonnin

Send message
Joined: 2 Jul 16
Posts: 339
Credit: 8,281,341,558
RAC: 2,803
Level
Tyr
Scientific publications
watwatwatwatwat
Message 48588 - Posted: 4 Jan 2018, 0:14:48 UTC - in response to Message 48587.  

50,000 WUs... Holy ****. If only those were GPU WUs


I saw that too and downloaded 4 on a computer I haven't run any of these on. Started at the same time and boom all went to crap. I even tried to pause some to stop the error.

http://www.gpugrid.net/results.php?hostid=458003
ID: 48588 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
STARBASEn
Avatar

Send message
Joined: 17 Feb 09
Posts: 91
Credit: 1,603,303,394
RAC: 0
Level
His
Scientific publications
watwatwatwatwatwatwatwatwatwat
Message 48589 - Posted: 4 Jan 2018, 1:02:26 UTC
Last modified: 4 Jan 2018, 1:05:56 UTC

So far from what I have learned starting more than one multiple cpu job at a time in a split core senerio is that they need not be started at exactly the same time for one to error. Given not an exact start at the same time, the one started first always errors and the second one started processes to successful completion with the next WU in the queue started to also complete successfully. Four of four tries, near simultaneous starts the one started first ended up failing. Secondly, controlling the time to when the second WU is allowed to start following the first start time is up to 5 seconds as tested so far. Third, when the boinc client switches between projects, the QC WU's so far observed are completed in pairs leaving no single job left in progress (suspended) to stagger start the times. This unfortunate characteristic means "simultaneous (or nearly so) starts" cause an error whenever the client switches back to the gpugrid cpu jobs with a queue larger than one WU. Guess the only way to prevent this behavior is to not split cores, especially for unattended clients.

Edit: correct my lousy spelling, stupid keyboard :)
ID: 48589 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Keith Myers
Avatar

Send message
Joined: 13 Dec 17
Posts: 1424
Credit: 9,189,946,190
RAC: 0
Level
Tyr
Scientific publications
watwatwatwatwat
Message 48590 - Posted: 4 Jan 2018, 1:15:38 UTC - in response to Message 48589.  

I didn't have that experience with the two TONI tasks I started simultaneously. Or within the 5 second window you described. Both completed successfully. I am limiting core usage to four with an app_config file. Limiting the max_concurrent to 1 now since I also crunch SETI cpu tasks on that computer. I ran the two concurrent jobs when Toni requested users to try that experiment.
ID: 48590 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Keith Myers
Avatar

Send message
Joined: 13 Dec 17
Posts: 1424
Credit: 9,189,946,190
RAC: 0
Level
Tyr
Scientific publications
watwatwatwatwat
Message 48591 - Posted: 4 Jan 2018, 2:02:50 UTC

Wow, the credit awarded is all over the place for these DOMINIK tasks. Obviously NOT tied to computation time or resources used for compute.

Task 16862848
3715 seconds CPU time Credit awarded 21

Task 16815159
3687 seconds CPU time Credit awarded 161
ID: 48591 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
FredoGuan

Send message
Joined: 29 Dec 16
Posts: 2
Credit: 1,397
RAC: 0
Level

Scientific publications
wat
Message 48592 - Posted: 4 Jan 2018, 2:12:21 UTC

I just got a task and it finished on a dual e5-2450l 32g ram server fine.
resultid=16815237
Keep developing this, please. This is quite nice and I would really like to see this as part of GPUGRID permanently.
ID: 48592 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
FredoGuan

Send message
Joined: 29 Dec 16
Posts: 2
Credit: 1,397
RAC: 0
Level

Scientific publications
wat
Message 48593 - Posted: 4 Jan 2018, 2:57:04 UTC - in response to Message 48592.  

resultid=16815470
ID: 48593 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Dominik

Send message
Joined: 15 Dec 17
Posts: 9
Credit: 0
RAC: 0
Level

Scientific publications
wat
Message 48597 - Posted: 4 Jan 2018, 11:10:04 UTC

Hello Keith,

Wow, the credit awarded is all over the place for these DOMINIK tasks. Obviously NOT tied to computation time or resources used for compute.


Did you observe this behavior multiple times? Really strange to be honest.


Thanks for helping out everyone!
ID: 48597 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
mmonnin

Send message
Joined: 2 Jul 16
Posts: 339
Credit: 8,281,341,558
RAC: 2,803
Level
Tyr
Scientific publications
watwatwatwatwat
Message 48603 - Posted: 4 Jan 2018, 13:07:24 UTC - in response to Message 48597.  

All tasks were erroring out in 2 minutes due to the app using gcc5.5

This got it to go to farther and start being multithreaded. We'll see if it actually completes.

sudo apt-get install gcc-5 g++-5
ID: 48603 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
mmonnin

Send message
Joined: 2 Jul 16
Posts: 339
Credit: 8,281,341,558
RAC: 2,803
Level
Tyr
Scientific publications
watwatwatwatwat
Message 48608 - Posted: 4 Jan 2018, 14:08:09 UTC - in response to Message 48603.  

Yup, it completed.
http://www.gpugrid.net/result.php?resultid=16817115
ID: 48608 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Dominik

Send message
Joined: 15 Dec 17
Posts: 9
Credit: 0
RAC: 0
Level

Scientific publications
wat
Message 48610 - Posted: 4 Jan 2018, 14:14:33 UTC

Great! Thank you very much
ID: 48610 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
klepel

Send message
Joined: 23 Dec 09
Posts: 189
Credit: 4,813,881,008
RAC: 149
Level
Arg
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 48613 - Posted: 4 Jan 2018, 16:01:40 UTC - in response to Message 48401.  

@klepel - can you try installing gcc (if not already there)?

tks

I tried it yesterday. I installed gcc-5 and gcc-6. And it worked on the computer http://www.gpugrid.net/results.php?hostid=452211
ID: 48613 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
STARBASEn
Avatar

Send message
Joined: 17 Feb 09
Posts: 91
Credit: 1,603,303,394
RAC: 0
Level
His
Scientific publications
watwatwatwatwatwatwatwatwatwat
Message 48616 - Posted: 4 Jan 2018, 16:49:17 UTC - in response to Message 48590.  

These are the two WU's that were started about 5 seconds apart on an AMD FX-8350 with the first one started failing.

Stdoutdea.txt:
03-Jan-2018 17:20:14 [GPUGRID] [css] running e113s22_e86s4p0f123-PABLO_p53_PHEX10P_IDP-0-1-RND2720_0 (0.987 CPUs + 1 NVIDIA GPU)
03-Jan-2018 17:20:14 [GPUGRID] Starting task c00000_00024-DOMINIK_QMML2_m0000000055-0-1-RND3244_0
03-Jan-2018 17:20:14 [GPUGRID] [cpu_sched] Starting task c00000_00024-DOMINIK_QMML2_m0000000055-0-1-RND3244_0 using QC version 314 (mt) in slot 9
03-Jan-2018 17:20:14 [GPUGRID] [css] running c00000_00024-DOMINIK_QMML2_m0000000055-0-1-RND3244_0 (4 CPUs)
03-Jan-2018 17:20:20 [GPUGRID] task c06475_06499-DOMINIK_QMML2_m0000000054-0-1-RND3067_0 resumed by user
03-Jan-2018 17:20:21 [GPUGRID] [css] running e113s22_e86s4p0f123-PABLO_p53_PHEX10P_IDP-0-1-RND2720_0 (0.987 CPUs + 1 NVIDIA GPU)
03-Jan-2018 17:20:21 [GPUGRID] [css] running c00000_00024-DOMINIK_QMML2_m0000000055-0-1-RND3244_0 (4 CPUs)
03-Jan-2018 17:20:21 [GPUGRID] Starting task c06475_06499-DOMINIK_QMML2_m0000000054-0-1-RND3067_0
03-Jan-2018 17:20:21 [GPUGRID] [cpu_sched] Starting task c06475_06499-DOMINIK_QMML2_m0000000054-0-1-RND3067_0 using QC version 314 (mt) in slot 10
03-Jan-2018 17:20:21 [GPUGRID] [css] running c06475_06499-DOMINIK_QMML2_m0000000054-0-1-RND3067_0 (4 CPUs)
03-Jan-2018 17:20:25 [GPUGRID] [sched_op] Deferring communication for 00:01:39
03-Jan-2018 17:20:25 [GPUGRID] [sched_op] Reason: Unrecoverable error for task c00000_00024-DOMINIK_QMML2_m0000000055-0-1-RND3244_0
03-Jan-2018 17:20:25 [GPUGRID] Computation for task c00000_00024-DOMINIK_QMML2_m0000000055-0-1-RND3244_0 finished
03-Jan-2018 17:20:25 [GPUGRID] [css] running e113s22_e86s4p0f123-PABLO_p53_PHEX10P_IDP-0-1-RND2720_0 (0.987 CPUs + 1 NVIDIA GPU)
03-Jan-2018 17:20:25 [GPUGRID] [css] running c06475_06499-DOMINIK_QMML2_m0000000054-0-1-RND3067_0 (4 CPUs)


I didn't have that experience with the two TONI tasks I started simultaneously. Or within the 5 second window you described. Both completed successfully. I am limiting core usage to four with an app_config file. Limiting the max_concurrent to 1 now since I also crunch SETI cpu tasks on that computer. I ran the two concurrent jobs when Toni requested users to try that experiment.


Since this hasn't been the case with your Intel's implies this could be a cpu related phenomenom (architecture/scheduling differences). Perhaps the Intel's can handle initial start up processes faster than the FX series AMD, (might spring for a Ryen7 soon just to check them as well). Regardless, the issue is resolved with my systems by limiting concurrent QC jobs to one and use the other four cores to run WCG as to date I have not experienced a concurrent issue with the WCG WU's.
ID: 48616 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Keith Myers
Avatar

Send message
Joined: 13 Dec 17
Posts: 1424
Credit: 9,189,946,190
RAC: 0
Level
Tyr
Scientific publications
watwatwatwatwat
Message 48617 - Posted: 4 Jan 2018, 18:14:01 UTC - in response to Message 48597.  

Hello Keith,

Wow, the credit awarded is all over the place for these DOMINIK tasks. Obviously NOT tied to computation time or resources used for compute.


Did you observe this behavior multiple times? Really strange to be honest.


Thanks for helping out everyone!

Yes, the first completed tasks got reasonable credit. Then when I downloaded more, all the credit for them nosedived. Once I saw that they weren't worth running I set NNT.




    16862848 12962574 456812 3 Jan 2018 | 23:38:10 UTC 4 Jan 2018 | 1:20:46 UTC Completed and validated 1,000.69 3,715.60 21.12 Quantum Chemistry v3.14 (mt)
    16815333 12963093 456812 4 Jan 2018 | 1:00:00 UTC 4 Jan 2018 | 2:57:45 UTC Completed and validated 1,020.00 3,812.63 27.38 Quantum Chemistry v3.14 (mt)
    16815332 12963092 456812 4 Jan 2018 | 0:59:23 UTC 4 Jan 2018 | 2:40:47 UTC Completed and validated 990.38 3,697.44 25.78 Quantum Chemistry v3.14 (mt)
    16815320 12963080 456812 4 Jan 2018 | 1:00:37 UTC 4 Jan 2018 | 3:14:28 UTC Completed and validated 997.66 3,731.61 27.28 Quantum Chemistry v3.14 (mt)
    16815307 12963067 456812 4 Jan 2018 | 1:07:01 UTC 4 Jan 2018 | 4:04:01 UTC Completed and validated 1,009.28 3,642.48 26.40 Quantum Chemistry v3.14 (mt)
    16815275 12963035 456812 4 Jan 2018 | 1:07:38 UTC 4 Jan 2018 | 4:21:09 UTC Completed and validated 1,033.63 3,668.31 26.52 Quantum Chemistry v3.14 (mt)
    16815264 12963024 456812 4 Jan 2018 | 1:06:24 UTC 4 Jan 2018 | 3:46:54 UTC Completed and validated 969.15 3,592.69 25.72 Quantum Chemistry v3.14 (mt)
    16815248 12963008 456812 4 Jan 2018 | 0:58:46 UTC 4 Jan 2018 | 2:24:17 UTC Completed and validated 935.31 3,503.52 23.59 Quantum Chemistry v3.14 (mt)
    16815234 12962994 456812 4 Jan 2018 | 1:05:48 UTC 4 Jan 2018 | 3:30:44 UTC Completed and validated 981.30 3,616.45 26.43 Quantum Chemistry v3.14 (mt)
    16815171 12962931 456812 3 Jan 2018 | 23:37:35 UTC 4 Jan 2018 | 1:04:10 UTC Completed and validated 929.82 3,523.44 18.45 Quantum Chemistry v3.14 (mt)
    16815159 12962919 456812 3 Jan 2018 | 23:35:43 UTC 4 Jan 2018 | 0:16:22 UTC Completed and validated 986.88 3,687.96 161.36 Quantum Chemistry v3.14 (mt)

ID: 48617 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
klepel

Send message
Joined: 23 Dec 09
Posts: 189
Credit: 4,813,881,008
RAC: 149
Level
Arg
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 48618 - Posted: 4 Jan 2018, 19:31:13 UTC
Last modified: 4 Jan 2018, 19:36:13 UTC

I have to report back on the AMD Ryzen 1700x Computer: http://www.gpugrid.net/results.php?hostid=420971

If I run 3 instances, the WUs crashes after about 200 seconds, and after that the computer crashes completely. If I am running one (01) instance (WU), the computer runs without any problem.

However, as BOINC downloads several of this Quantum Chemistry v3.14 (mt) WUs, BOINC thinks my CPU cache is full and refuses to download additional CPU WUs from PRIMGRID. So after a while the CPU is only loaded with one QC WU (4 threads) and the rest of the cores are idle - Not very efficient.

Sorry, Dominik unter diesen Umständen kann ich keine weiteren QC WUs für diesen Computer herunterladen. Komme aber gerne zurück, wenn wir ohne Probleme mehrere MultiCores WUs gleichzeitig bearbeiten können.
ID: 48618 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
STARBASEn
Avatar

Send message
Joined: 17 Feb 09
Posts: 91
Credit: 1,603,303,394
RAC: 0
Level
His
Scientific publications
watwatwatwatwatwatwatwatwatwat
Message 48619 - Posted: 4 Jan 2018, 19:45:44 UTC - in response to Message 48618.  
Last modified: 4 Jan 2018, 20:02:00 UTC

However, as BOINC downloads several of this Quantum Chemistry v3.14 (mt) WUs, BOINC thinks my CPU cache is full and refuses to download additional CPU WUs from PRIMGRID. So after a while the CPU is only loaded with one QC WU (4 threads) and the rest of the cores are idle - Not very efficient.


If you temporarily suspend the QC jobs (except perhaps the one in progress), you should download more work from your other projects and once downloaded, resume the QC jobs and let boinc take over running the various projects as you have them configured. You may need to "update" the projects you want more work from under the "projects" tab to initiate the downloads right away.

Edit: Close quote and add last sentence.
ID: 48619 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Keith Myers
Avatar

Send message
Joined: 13 Dec 17
Posts: 1424
Credit: 9,189,946,190
RAC: 0
Level
Tyr
Scientific publications
watwatwatwatwat
Message 48624 - Posted: 5 Jan 2018, 2:50:48 UTC

Seti@home has been down all day so ran out of work. Decided to give the QC Dominik tasks another try. Thought that possibly the last batch were unique or one-offs or something different about the computer from when I first ran them.

Nope. Even worse credit for the batch I ran this afternoon. Credits awarded = 6.

If you expect to get anybody to want to run these, you are going to have to make them more appealing, credit-wise at least. For me, not worth the electricity to run them. Would rather let the computer go cold and give my power bill a temporary reprieve.
ID: 48624 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
mmonnin

Send message
Joined: 2 Jul 16
Posts: 339
Credit: 8,281,341,558
RAC: 2,803
Level
Tyr
Scientific publications
watwatwatwatwat
Message 48632 - Posted: 5 Jan 2018, 11:58:50 UTC
Last modified: 5 Jan 2018, 12:07:27 UTC

Yeah credit took a dump on the last two I completed.

Run Time-----CPU Time-----Credit
2,389.43-----26,206.90-----549.62
3,034.92-----36,710.33-----94.32
ID: 48632 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
bormolino

Send message
Joined: 16 May 13
Posts: 41
Credit: 145,731,947
RAC: 0
Level
Cys
Scientific publications
watwatwatwatwatwat
Message 48639 - Posted: 5 Jan 2018, 20:01:49 UTC

I got 5.97 points for 32 minutes calculation on my fx 6100.

That's ridiculous!
ID: 48639 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Keith Myers
Avatar

Send message
Joined: 13 Dec 17
Posts: 1424
Credit: 9,189,946,190
RAC: 0
Level
Tyr
Scientific publications
watwatwatwatwat
Message 48642 - Posted: 5 Jan 2018, 21:16:29 UTC

If you care to learn about why the low credit or why certain tasks get hi-middle-low credit, read my post here
ID: 48642 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Previous · 1 . . . 3 · 4 · 5 · 6 · 7 · Next

Message boards : Multicore CPUs : New batch of QC tasks (QMML)

©2026 Universitat Pompeu Fabra