Simultaneously starting MCs

Message boards : Multicore CPUs : Simultaneously starting MCs
Message board moderation

To post messages, you must log in.

1 · 2 · 3 · Next

AuthorMessage
DRSMT

Send message
Joined: 23 Feb 17
Posts: 21
Credit: 5,528,199,475
RAC: 0
Level
Tyr
Scientific publications
watwatwatwat
Message 49378 - Posted: 2 May 2018, 11:38:08 UTC

Problem with simultaneously starting Multicore CPU tasks has not been fixed yet!
ID: 49378 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
PappaLitto

Send message
Joined: 21 Mar 16
Posts: 513
Credit: 4,673,458,277
RAC: 0
Level
Arg
Scientific publications
watwatwatwatwatwatwatwat
Message 49382 - Posted: 2 May 2018, 14:48:28 UTC - in response to Message 49378.  

Problem with simultaneously starting Multicore CPU tasks has not been fixed yet!

I have heard about this error but maybe I don't understand the symptoms. I have had three 4 thread WUs start at once and sometimes they work, sometimes they don't. Could this be the cause of my errors? Linked below are the tasks from the system:

http://www.gpugrid.net/results.php?hostid=424454
ID: 49382 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
DRSMT

Send message
Joined: 23 Feb 17
Posts: 21
Credit: 5,528,199,475
RAC: 0
Level
Tyr
Scientific publications
watwatwatwat
Message 49383 - Posted: 2 May 2018, 16:15:15 UTC - in response to Message 49382.  

If two WUs start at the same time, in most cases one of the WU failes right at start and throws an calculation error, which is very inconveniant, if you have to start often times more than one WU at the same time (in my case up to 20 WUs on my 80 threads machine). Would like to hear some statement of the developer(s) or simply a bugfix within the WUs, because at the moment, I do not see any suitable work around. Don't think it's just your fault or mine, because several users have reported this issue for a while now.
ID: 49383 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Toni
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist

Send message
Joined: 9 Dec 08
Posts: 1006
Credit: 5,068,599
RAC: 0
Level
Ser
Scientific publications
watwatwatwat
Message 49614 - Posted: 6 Jun 2018, 12:29:25 UTC - in response to Message 49383.  
Last modified: 6 Jun 2018, 12:34:48 UTC

If it is still failing, please provide a task number for me to check.

@Thomas: pls try to reset the project
ID: 49614 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
captainjack

Send message
Joined: 9 May 13
Posts: 171
Credit: 4,739,796,466
RAC: 1,441
Level
Arg
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 49615 - Posted: 6 Jun 2018, 12:32:41 UTC
Last modified: 6 Jun 2018, 12:36:41 UTC

Toni,

Here are two task numbers that started at the same time. Both failed.

17737243
17737150

Let me know if you need more info.

EDIT:

Here is a task that started by itself and failed.

17714391
ID: 49615 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Toni
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist

Send message
Joined: 9 Dec 08
Posts: 1006
Credit: 5,068,599
RAC: 0
Level
Ser
Scientific publications
watwatwatwat
Message 49616 - Posted: 6 Jun 2018, 12:36:17 UTC - in response to Message 49615.  
Last modified: 6 Jun 2018, 12:37:19 UTC

@captainjack: please try two things for me

1. open a terminal, and run the
flock
command. See if it gives an error (command not found) or a longer message.
2. reset the project

Thanks
ID: 49616 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
captainjack

Send message
Joined: 9 May 13
Posts: 171
Credit: 4,739,796,466
RAC: 1,441
Level
Arg
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 49617 - Posted: 6 Jun 2018, 12:50:46 UTC

Toni,

The "flock" command was found and asked for more arguments.

After a project reset, the following two tasks were started and both failed.

17737309
17737152

The following task was started by itself and failed.

17737252

What next?
ID: 49617 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Toni
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist

Send message
Joined: 9 Dec 08
Posts: 1006
Credit: 5,068,599
RAC: 0
Level
Ser
Scientific publications
watwatwatwat
Message 49618 - Posted: 6 Jun 2018, 12:55:33 UTC - in response to Message 49617.  

:(

Which means that for your host the new app is a regression. I need an enlightenment.

T
ID: 49618 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Stefan
Project administrator
Project developer
Project tester
Project scientist

Send message
Joined: 5 Mar 13
Posts: 348
Credit: 0
RAC: 0
Level

Scientific publications
wat
Message 49620 - Posted: 6 Jun 2018, 13:18:42 UTC

Statistically though the new app seems to have worked on other hosts. We went from 900 WU to 1500 WU in progress.
ID: 49620 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
DRSMT

Send message
Joined: 23 Feb 17
Posts: 21
Credit: 5,528,199,475
RAC: 0
Level
Tyr
Scientific publications
watwatwatwat
Message 49621 - Posted: 6 Jun 2018, 13:20:38 UTC

with all my computers just the same... Toni, does it help if I give you by private message the remote control access credentials of one of my Linux machines, so you can test on your own?
ID: 49621 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Toni
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist

Send message
Joined: 9 Dec 08
Posts: 1006
Credit: 5,068,599
RAC: 0
Level
Ser
Scientific publications
watwatwatwat
Message 49622 - Posted: 6 Jun 2018, 13:40:14 UTC - in response to Message 49621.  

Thomas, that would help, but perhaps let me ask another thing first:

What OS do you have, and which procedure did you use to install boinc?

(Also, I assume QM tasks were working ok before, right?)
ID: 49622 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
DRSMT

Send message
Joined: 23 Feb 17
Posts: 21
Credit: 5,528,199,475
RAC: 0
Level
Tyr
Scientific publications
watwatwatwat
Message 49623 - Posted: 6 Jun 2018, 13:52:34 UTC - in response to Message 49622.  

Sometimes they work and sometimes not. If two or more WUs start at the same time, they all throw calculation errors. This was the state until now. But with the very new version you just released today, it seems like they are not working anymore at all. My operating system is Linux Mint 18.3 64 Bit with actual linux kernel. I installed boinc with "sudo apt-get install boinc". gcc-5 and g++-5 are installed; also python-support.
ID: 49623 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Toni
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist

Send message
Joined: 9 Dec 08
Posts: 1006
Credit: 5,068,599
RAC: 0
Level
Ser
Scientific publications
watwatwatwat
Message 49624 - Posted: 6 Jun 2018, 14:16:27 UTC - in response to Message 49623.  
Last modified: 6 Jun 2018, 14:18:08 UTC

Ok thanks. This will need a while to debug. As you can see there is no error info to help (as usual in boinc...).
ID: 49624 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
[VENETO] boboviz

Send message
Joined: 10 Sep 10
Posts: 164
Credit: 388,132
RAC: 0
Level

Scientific publications
wat
Message 49625 - Posted: 6 Jun 2018, 14:33:24 UTC
Last modified: 6 Jun 2018, 14:43:00 UTC

All errors on my VM (with 4 virtual core).

<message>
process exited with code 195 (0xc3, -61)
</message>
<stderr_txt>
16:29:25 (2940): wrapper (7.7.26016): starting
16:29:25 (2940): wrapper (7.7.26016): starting
16:29:25 (2940): wrapper: running /bin/bash (-c "flock /var/lib/boinc-client/projects/www.gpugrid.net/miniconda.lock ./miniconda-installer -b -u -p /var/lib/boinc-client/projects/www.gpugrid.net/miniconda")
Please run using "bash" or "sh", but not "." or "source"\n16:29:26 (2940): /bin/bash exited; CPU time 0.000000
16:29:26 (2940): app exit status: 0x1


I have not app_config.

Addendum: i also tried to stop all wus and started manually one-by-one. Same error.
ID: 49625 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Toni
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist

Send message
Joined: 9 Dec 08
Posts: 1006
Credit: 5,068,599
RAC: 0
Level
Ser
Scientific publications
watwatwatwat
Message 49626 - Posted: 6 Jun 2018, 14:57:07 UTC - in response to Message 49624.  

Version 320 out
ID: 49626 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Toni
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist

Send message
Joined: 9 Dec 08
Posts: 1006
Credit: 5,068,599
RAC: 0
Level
Ser
Scientific publications
watwatwatwat
Message 49627 - Posted: 6 Jun 2018, 15:10:13 UTC - in response to Message 49626.  

Also, you need the libc6-dev package

sudo apt install libc6-dev
ID: 49627 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
bormolino

Send message
Joined: 16 May 13
Posts: 41
Credit: 145,731,947
RAC: 0
Level
Cys
Scientific publications
watwatwatwatwatwat
Message 49629 - Posted: 6 Jun 2018, 15:29:59 UTC

All tasks exit with error:

14:19:18 (8806): wrapper: running /bin/bash (-c "flock /var/lib/boinc-client/projects/www.gpugrid.net/miniconda.lock ./miniconda-installer -b -u -p /var/lib/boinc-client/projects/www.gpugrid.net/miniconda")
Please run using "bash" or "sh", but not "." or "source"\n14:19:19 (8806): /bin/bash exited; CPU time 0.001596
14:19:19 (8806): app exit status: 0x1
ID: 49629 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Toni
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist

Send message
Joined: 9 Dec 08
Posts: 1006
Credit: 5,068,599
RAC: 0
Level
Ser
Scientific publications
watwatwatwat
Message 49630 - Posted: 6 Jun 2018, 15:32:00 UTC - in response to Message 49629.  

These were 3.19. See with 3.20
ID: 49630 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
captainjack

Send message
Joined: 9 May 13
Posts: 171
Credit: 4,739,796,466
RAC: 1,441
Level
Arg
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 49631 - Posted: 6 Jun 2018, 15:51:43 UTC

When I tried to start two at the same time, one of them runs okay and the other one aborts.

Task id for the aborted work unit is 17714759
Work unit number for the aborted work unit is 13679491

Running version 3.20

Let me know if you need more info.
ID: 49631 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
DRSMT

Send message
Joined: 23 Feb 17
Posts: 21
Credit: 5,528,199,475
RAC: 0
Level
Tyr
Scientific publications
watwatwatwat
Message 49632 - Posted: 6 Jun 2018, 17:57:54 UTC

I have had libc6-dev already installed... Earlier I got a bunch of new WUs, but all failed after several minutes of calculation (~ 5 - 15 minutes).
ID: 49632 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
1 · 2 · 3 · Next

Message boards : Multicore CPUs : Simultaneously starting MCs

©2026 Universitat Pompeu Fabra