New QC app

Message boards : Multicore CPUs : New QC app
Message board moderation

To post messages, you must log in.

1 · 2 · 3 · 4 . . . 7 · Next

AuthorMessage
Toni
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist

Send message
Joined: 9 Dec 08
Posts: 1006
Credit: 5,068,599
RAC: 0
Level
Ser
Scientific publications
watwatwatwat
Message 49759 - Posted: 2 Jul 2018, 10:18:54 UTC
Last modified: 2 Jul 2018, 10:52:36 UTC

Dears, after a hot weekend during which I accidentally cancelled QC WUs, we are ready to start again with a new app. As soon as we get it right, we should be able to run on more machines (gcc no longer a requirement).

There will be a largish download the first time you run app 329. If you want to free up some disk space, please reset the project (recommended, but not urgent).
ID: 49759 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
kain

Send message
Joined: 3 Sep 14
Posts: 152
Credit: 927,557,369
RAC: 108
Level
Glu
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwat
Message 49760 - Posted: 2 Jul 2018, 11:29:25 UTC

Ready for action ;)
ID: 49760 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
tullio

Send message
Joined: 8 May 18
Posts: 190
Credit: 104,426,808
RAC: 0
Level
Cys
Scientific publications
wat
Message 49766 - Posted: 2 Jul 2018, 16:30:41 UTC

First 330 task completed and validated. The second one is waiting for memory alongside a GPU task, which is however using only 4% of the total 8 GB of RAM.
Tullio
ID: 49766 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Stefan
Project administrator
Project developer
Project tester
Project scientist

Send message
Joined: 5 Mar 13
Posts: 348
Credit: 0
RAC: 0
Level

Scientific publications
wat
Message 49778 - Posted: 4 Jul 2018, 8:47:07 UTC - in response to Message 49766.  

I will let the WUs run out for a day because I want to see if something weird is happening on my side (the WUs are calculating fine, don't worry). I'll submit more once they are completed tomorrow.
ID: 49778 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Richard Haselgrove

Send message
Joined: 11 Jul 09
Posts: 1639
Credit: 10,159,968,649
RAC: 0
Level
Trp
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 50348 - Posted: 30 Aug 2018, 13:43:18 UTC

@ Toni, @ Stefan

There's an error report in Number Crunching (This computer has finished a daily quota of 31 tasks) which suggests that the maximum upload size for a batch of QC tasks has been set too low.

Task name is 6955_1_15_16_18_dd130713_n00001-SDOERR_SELE2-0-1-RND2528
ID: 50348 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Richard Haselgrove

Send message
Joined: 11 Jul 09
Posts: 1639
Credit: 10,159,968,649
RAC: 0
Level
Trp
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 50349 - Posted: 30 Aug 2018, 13:43:42 UTC

@ Toni, @ Stefan

There's an error report in Number Crunching (This computer has finished a daily quota of 31 tasks) which suggests that the maximum upload size for a batch of QC tasks has been set too low.

Task name is 6955_1_15_16_18_dd130713_n00001-SDOERR_SELE2-0-1-RND2528
ID: 50349 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Toni
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist

Send message
Joined: 9 Dec 08
Posts: 1006
Credit: 5,068,599
RAC: 0
Level
Ser
Scientific publications
watwatwatwat
Message 50404 - Posted: 5 Sep 2018, 14:34:10 UTC - in response to Message 50349.  
Last modified: 5 Sep 2018, 14:34:49 UTC

I think the actual error is a segmentation fault, which leaves large temporary files behind, and their transfer is attempted. Let's see if the situation improves with the new version.

If in doubt, please reset the project.
ID: 50404 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Zalster
Avatar

Send message
Joined: 26 Feb 14
Posts: 211
Credit: 4,496,324,562
RAC: 0
Level
Arg
Scientific publications
watwatwatwatwatwatwatwat
Message 50412 - Posted: 6 Sep 2018, 5:32:45 UTC - in response to Message 50404.  
Last modified: 6 Sep 2018, 5:33:01 UTC

Just curious, is there someplace that tells how well the QC apps are running like on the server status page. I can look at my own machines and see the errors there but overall is there someplace like for the GPU apps?
ID: 50412 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Richard Haselgrove

Send message
Joined: 11 Jul 09
Posts: 1639
Credit: 10,159,968,649
RAC: 0
Level
Trp
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 50415 - Posted: 6 Sep 2018, 8:25:58 UTC - in response to Message 50404.  

I think the actual error is a segmentation fault, which leaves large temporary files behind, and their transfer is attempted. Let's see if the situation improves with the new version.

BOINC won't attempt to upload a temporary file unless its name is specified with an upload URL in the workunit template.

I'll be able to advise better when you release the Windows app, and I can see any problems happening on my own machines.
ID: 50415 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Toni
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist

Send message
Joined: 9 Dec 08
Posts: 1006
Credit: 5,068,599
RAC: 0
Level
Ser
Scientific publications
watwatwatwat
Message 50417 - Posted: 6 Sep 2018, 8:46:57 UTC - in response to Message 50415.  

May well be restricted to a few machines.
ID: 50417 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Richard Haselgrove

Send message
Joined: 11 Jul 09
Posts: 1639
Credit: 10,159,968,649
RAC: 0
Level
Trp
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 50418 - Posted: 6 Sep 2018, 9:42:25 UTC - in response to Message 50417.  

May well be restricted to a few machines.

If you want to count me in, I'll do my best to report on any issues that may arise (in beta mode if necessary). I'm primarily Windows 7, so the WSL approach would be difficult except for one dual-boot test machine with Windows 10 ready to run.
ID: 50418 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
tullio

Send message
Joined: 8 May 18
Posts: 190
Credit: 104,426,808
RAC: 0
Level
Cys
Scientific publications
wat
Message 50419 - Posted: 6 Sep 2018, 10:21:20 UTC

My main Linux box is crunching SELE6 on its Opteron 1210. If necessary, I have a Windows 10 PC with an AMD A10-6700 and 22 GB RAM.
Tullio
ID: 50419 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
[VENETO] boboviz

Send message
Joined: 10 Sep 10
Posts: 164
Credit: 388,132
RAC: 0
Level

Scientific publications
wat
Message 50420 - Posted: 6 Sep 2018, 12:24:27 UTC

again
<message>
upload failure: <file_xfer_error>
<file_name>5516_14_15_18_19_8125b500_n00001-SDOERR_SELE6-0-1-RND8343_0_1</file_name>
<error_code>-131 (file size too big)</error_code>
</file_xfer_error>


Have i had to reset the project?
ID: 50420 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Profile Chilean
Avatar

Send message
Joined: 8 Oct 12
Posts: 98
Credit: 385,652,461
RAC: 0
Level
Asp
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 50421 - Posted: 6 Sep 2018, 12:26:27 UTC

CPU apps not working very well. Lots of idle time. It maybe because of the 48 threads... some pythons use 4 cores, the rest just 1 and not all the time. It's gotta be an I/O issue (?).
ID: 50421 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Toni
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist

Send message
Joined: 9 Dec 08
Posts: 1006
Credit: 5,068,599
RAC: 0
Level
Ser
Scientific publications
watwatwatwat
Message 50422 - Posted: 6 Sep 2018, 13:18:05 UTC - in response to Message 50421.  

Each task should use 4 threads max.
ID: 50422 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
[VENETO] boboviz

Send message
Joined: 10 Sep 10
Posts: 164
Credit: 388,132
RAC: 0
Level

Scientific publications
wat
Message 50423 - Posted: 6 Sep 2018, 15:16:03 UTC - in response to Message 50420.  

again
<message>
upload failure: <file_xfer_error>
<file_name>5516_14_15_18_19_8125b500_n00001-SDOERR_SELE6-0-1-RND8343_0_1</file_name>
<error_code>-131 (file size too big)</error_code>
</file_xfer_error>


Have i had to reset the project?


Another clue: this happens when i reboot the virtual machine (and the wu restarts)
ID: 50423 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Jim1348

Send message
Joined: 28 Jul 12
Posts: 819
Credit: 1,591,285,971
RAC: 0
Level
His
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 50424 - Posted: 6 Sep 2018, 17:11:33 UTC
Last modified: 6 Sep 2018, 17:13:33 UTC

I have had good luck for the past day running QC on my i7-4770 (Ubuntu 16.04). That doesn't prove much, except that there is no fatal flaw in all the work units.

And I am limiting them to two cores per work unit, and three work units at a time. That gives me essentially the same output as four cores on two work units at a time, but leaves over a little more CPU support for my GTX 1070 on Folding. All in all, it seems to be working fine.

(boboviz - I wouldn't draw conclusions from virtual machines. Even LHC has a hard time with their own stuff.)
ID: 50424 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Richard Haselgrove

Send message
Joined: 11 Jul 09
Posts: 1639
Credit: 10,159,968,649
RAC: 0
Level
Trp
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 50425 - Posted: 6 Sep 2018, 17:32:31 UTC - in response to Message 50423.  

(and the wu re-starts)

I think that is indeed a clue. One of the mechanisms I considered was a WU re-starting, and appending a second result to an existing file, doubling its size.

Checking the 'headroom' between the typical result file size and <max_nbytes> was one of the tests I had in mind for the Windows version. Can anyone comment?
ID: 50425 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
tullio

Send message
Joined: 8 May 18
Posts: 190
Credit: 104,426,808
RAC: 0
Level
Cys
Scientific publications
wat
Message 50426 - Posted: 6 Sep 2018, 18:15:16 UTC

Most LHC users are Windows users and they run Scientific Linux research programs from CERN (not BOINC programs) using Virtual Machines.
Tullio
ID: 50426 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Jim1348

Send message
Joined: 28 Jul 12
Posts: 819
Credit: 1,591,285,971
RAC: 0
Level
His
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 50427 - Posted: 7 Sep 2018, 11:41:06 UTC
Last modified: 7 Sep 2018, 11:42:54 UTC

This is interesting. Twice overnight my PC crashed. Shut down. Didn't run.
But each time, I saw no errors in the BoincTasks log, or in the Folding log either. The Folding work unit just continued from where it left off.

But each time, a new set of (three) QC tasks started after starting up the PC.

And now I see the same old error message in the stderr.txt file:

</stderr_txt>
<message>
upload failure: <file_xfer_error>
<file_name>2360_16_18_20_21_e0c95459_n00001-SDOERR_SELE6-0-1-RND3935_0_1</file_name>
<error_code>-131 (file size too big)</error_code>
</file_xfer_error>
</message>

http://www.gpugrid.net/result.php?resultid=18686175

The messages don't appear immediately after the crashes, but a few hours later.
And I just attached this machine to GPUGrid a couple of days ago. So it didn't take long.
ID: 50427 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
1 · 2 · 3 · 4 . . . 7 · Next

Message boards : Multicore CPUs : New QC app

©2026 Universitat Pompeu Fabra