New batch of QC tasks (QMML)

Message boards : Multicore CPUs : New batch of QC tasks (QMML)
Message board moderation

To post messages, you must log in.

1 · 2 · 3 · 4 . . . 7 · Next

AuthorMessage
Toni
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist

Send message
Joined: 9 Dec 08
Posts: 1006
Credit: 5,068,599
RAC: 0
Level
Ser
Scientific publications
watwatwatwat
Message 48356 - Posted: 13 Dec 2017, 17:30:24 UTC

These are called QMML, and rather experimental (more dependencies). Let's see how they work.
ID: 48356 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
captainjack

Send message
Joined: 9 May 13
Posts: 171
Credit: 4,739,796,466
RAC: 1,182
Level
Arg
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 48358 - Posted: 13 Dec 2017, 21:21:51 UTC

Toni,

I have one of these that has been running for over 2 hours and so far looks like it is only using one CPU (thread). It has 4 threads allocated to each task.

There are also a number of warnings messages in the stderr.txt that look like this:

/var/lib/boinc-client/projects/www.gpugrid.net/miniconda/envs/qmml/lib/python3.6/site-packages/tables/path.py:112: NaturalNameWarning: object name is not a valid Python identifier: '122'; it does not match the pattern ``^[a-zA-Z_][a-zA-Z0-9_]*$``; you will not be able to use natural naming to access this object; using ``getattr()`` will still work, though
NaturalNameWarning)


You should be able to see all of them when the task uploads.

Let me know if you need more info.
ID: 48358 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Toni
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist

Send message
Joined: 9 Dec 08
Posts: 1006
Credit: 5,068,599
RAC: 0
Level
Ser
Scientific publications
watwatwatwat
Message 48359 - Posted: 13 Dec 2017, 22:10:31 UTC - in response to Message 48358.  
Last modified: 13 Dec 2017, 22:14:22 UTC

Thanks, the warnings are expected and harmless.

The thread allocation has some bug. I would have expected to use more threads than allocated, not less, but hey. I'll be debugging.

Also I hope suspend/resume and the progress bar work (more or less), unlike the old "plain" QC tasks.
ID: 48359 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
STARBASEn
Avatar

Send message
Joined: 17 Feb 09
Posts: 91
Credit: 1,603,303,394
RAC: 0
Level
His
Scientific publications
watwatwatwatwatwatwatwatwatwat
Message 48362 - Posted: 14 Dec 2017, 0:44:20 UTC

Subj: Observation of QC CPU WU's Core Utilization

Regarding the last go around several weeks ago with the QC CPU WU's that ran successfully on my 8 and 4 core cpu's, the following was observed.

4-core (Phenom II 3 GHz) cpu: core utilization was consistantly near 100% all 4 cores. Work completed in about 40 minutes per WU calender time.

8-core (FX-8350 4 GHz) cpu's (2): 4-cores (alternatively) were utilized at 100% with the remaining cores at much lesser utilization. Work completed with about 20 minutes real time per WU.

With the most recent QC WU's (12/13/2017), one FX-8350 errors out consistantly and will not run the WU's (maybe it requires software or drivers not currently installed) and the other FX-8350 runs them but at a substanially less core utilization rate than the prior WU batches. So far, the first WU in progress is currently less than 40% complete with about 2 hours calender time invested. Not sure if these recent WU's are of the same length as the earlier ones however.

I have uploaded photos of ksysguard graphic cpu utilization that can be observed if interested with addresses below.

Basically, as I am sure that this is a work in progress and bugs need to be resolved but I would conclude so far that these WU's process efficiently on a 4-core system but that all 8-cores should be fully utilized to make it worth while sacrificing 8-cores to a single WU when 8 WU's from other projects use all 8 at 100% thereby being much more efficient. Haven't tried the suspend/resume yet but will when the opprotunity is available. The previous poster appears correct re thread utilization being one issue.

Screenshots:
http://members.toast.net/obc/computing/grid_computing/images/QC_4-core_cpu.png
http://members.toast.net/obc/computing/grid_computing/images/QC_8-core_cpu.png
http://members.toast.net/obc/computing/grid_computing/images/QC-12-13-2017.png

(Sorry, crude way to present photos but wanted to get them up for anyone interested)
ID: 48362 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
mmonnin

Send message
Joined: 2 Jul 16
Posts: 339
Credit: 8,281,341,558
RAC: 2,803
Level
Tyr
Scientific publications
watwatwatwatwat
Message 48363 - Posted: 14 Dec 2017, 0:52:44 UTC
Last modified: 14 Dec 2017, 0:54:04 UTC

Mine is also just a bit less than 1 thread with 7 available for CPU usage. Good thing I have another client available to keep the CPU busy. If the progress bar is correct it will take about 9.5 hours on 1950x at 3.75 GHz.
ID: 48363 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Toni
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist

Send message
Joined: 9 Dec 08
Posts: 1006
Credit: 5,068,599
RAC: 0
Level
Ser
Scientific publications
watwatwatwat
Message 48364 - Posted: 14 Dec 2017, 8:44:09 UTC - in response to Message 48363.  

The next batch (QMML313a) should respect the number of threads requested by your client.
ID: 48364 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Profile Conan

Send message
Joined: 25 Mar 09
Posts: 25
Credit: 582,385
RAC: 0
Level
Gly
Scientific publications
wat
Message 48365 - Posted: 14 Dec 2017, 11:26:38 UTC

The Multiple threaded work units that were sent out last month worked fine for me with no issues.

These new ones however are all failing (so far on two different 64bit Linux machines) with this error

ERROR conda.core.link:_execute_actions(337): An error occurred while installing package 'psi4::gcc-5-5.2.0-1'.
LinkError: post-link script failed for package psi4::gcc-5-5.2.0-1
running your command again with `-v` will provide additional information
location of failed script: /home/Conan/BOINC/projects/www.gpugrid.net/miniconda/envs/qmml/bin/.gcc-5-post-link.sh

When checking this path I found that there is nothing in the /envs/ folder, which is probably where the job is failing.

Conan
ID: 48365 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
mmonnin

Send message
Joined: 2 Jul 16
Posts: 339
Credit: 8,281,341,558
RAC: 2,803
Level
Tyr
Scientific publications
watwatwatwatwat
Message 48366 - Posted: 14 Dec 2017, 11:58:44 UTC

Hmm 1st one completed for me and the 2nd one is at around 86%.
ID: 48366 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
captainjack

Send message
Joined: 9 May 13
Posts: 171
Credit: 4,739,796,466
RAC: 1,182
Level
Arg
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 48370 - Posted: 14 Dec 2017, 19:36:39 UTC

Toni said,

The next batch (QMML313a) should respect the number of threads requested by your client.


It looks like this one does respect the number of threads requested by my client. My app_config specifies 4 threads and it looks to be using 4 threads.

Let me know if you need more info.
ID: 48370 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Jim1348

Send message
Joined: 28 Jul 12
Posts: 819
Credit: 1,591,285,971
RAC: 0
Level
His
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 48371 - Posted: 14 Dec 2017, 19:56:33 UTC

I am not able to get QC on my Ryzen 1700 machine running Ubuntu 17.10. I just get a "No tasks are available for Quantum Chemistry" message when I request them. However, I am able to get QC on my i7 3770 machine running Ubuntu 16.04 (both machines have BOINC 7.8.3). Both machines are set to the same profile (work), so they should be treated identically.

But I see that some people with AMD machines get work. Is this a bug or a feature?
ID: 48371 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Toni
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist

Send message
Joined: 9 Dec 08
Posts: 1006
Credit: 5,068,599
RAC: 0
Level
Ser
Scientific publications
watwatwatwat
Message 48373 - Posted: 14 Dec 2017, 23:43:10 UTC - in response to Message 48371.  

The reason why otherwise similar machines get/do not get work completely baffles me. I don't think it's related to the maker of the CPU. Perhaps with the history of tasks/host reliability or somesuch. With this respects BOINC is of no help.
ID: 48373 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Profile bcavnaugh

Send message
Joined: 8 Nov 13
Posts: 56
Credit: 1,002,640,163
RAC: 0
Level
Met
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwat
Message 48374 - Posted: 15 Dec 2017, 0:11:25 UTC - in response to Message 48356.  
Last modified: 15 Dec 2017, 0:12:32 UTC

These are called QMML, and rather experimental (more dependencies). Let's see how they work.


I would really like to get some tasks but seems they are not being given out ATM

http://www.gpugrid.net/show_host_detail.php?hostid=457056
Been Trying for awhile now. Intel(R) Core(TM) i7-3970X

Crunching@EVGA The Number One Team in the BOINC Community. Folding@EVGA The Number One Team in the Folding@Home Community.
ID: 48374 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Jim1348

Send message
Joined: 28 Jul 12
Posts: 819
Credit: 1,591,285,971
RAC: 0
Level
His
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 48375 - Posted: 15 Dec 2017, 2:10:38 UTC - in response to Message 48373.  

The reason why otherwise similar machines get/do not get work completely baffles me.

I have seen many instances of it myself over the years, but had hoped the latest BOINC clients were past that. Unfortunately not.
ID: 48375 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
mmonnin

Send message
Joined: 2 Jul 16
Posts: 339
Credit: 8,281,341,558
RAC: 2,803
Level
Tyr
Scientific publications
watwatwatwatwat
Message 48376 - Posted: 15 Dec 2017, 2:54:01 UTC
Last modified: 15 Dec 2017, 3:00:20 UTC

My 1950x also on 17.10 is getting tasks.

Were the old 3.13 tasks producing bad data as my processing time was just wasted by the server canceling them in the middle of processing them? Thats the absolute worst thing a project admin can do. Cancel ones not started but don't ever take a crap on donated resources.

Its still not working right. In 22.5min of the task running it has used 1:37min of CPU time when the task is limited to 3 cores. That's over 4 cores of CPU usage. And they just had a computation error.

At least 3.13 worked.

http://www.gpugrid.net/result.php?resultid=16767178

Prob a good thing tasks are being sent to some AMD CPUs. Damn seg fault.
ID: 48376 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
klepel

Send message
Joined: 23 Dec 09
Posts: 189
Credit: 4,813,881,008
RAC: 149
Level
Arg
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 48378 - Posted: 15 Dec 2017, 4:25:33 UTC

Most of the Quantum Chemistry v3.14 (mt) fail on my AMD 1700x. v3.13 (mt) worked more or less.

As an example:
http://www.gpugrid.net/result.php?resultid=16767309

I use an app_config to limit the use to 4 cores for each work unit (WU) and runs 3 WUs in parallel. Two cores are reserved for the GPU.

I had to change the configuration to accept only GPU Work Unites as the computer crashed twice today.

Hope this helps.
ID: 48378 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Toni
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist

Send message
Joined: 9 Dec 08
Posts: 1006
Credit: 5,068,599
RAC: 0
Level
Ser
Scientific publications
watwatwatwat
Message 48379 - Posted: 15 Dec 2017, 8:58:27 UTC - in response to Message 48378.  
Last modified: 15 Dec 2017, 8:59:36 UTC

I understand that seing cancelled WUs is not nice, but it saves future crunching and network bandwidth (both server and client) that would otherwise be lost. Also, I thought that the function we use only cancelled UNSENT or un started wus.
ID: 48379 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
3de64piB5uZAS6SUNt1GFDU9dRhY
Avatar

Send message
Joined: 20 Apr 15
Posts: 285
Credit: 1,102,216,607
RAC: 0
Level
Met
Scientific publications
watwatwatwatwatwat
Message 48380 - Posted: 15 Dec 2017, 15:20:57 UTC
Last modified: 15 Dec 2017, 15:24:44 UTC

Hey friends, in case you need some more machines for testing, I can set up another one that is Linux based. As I have seen in the below conversation, there might be some issues with the CPU type. So .. I have both brands available for testing. Which one would help you most, Intel or AMD?

If you want me to, I could even let you choose the generation. From older Sandy/Ivy Bridge to new Skylake to Ryzen. Lust let me know and I will make one available on short notice.

Edit: I can even contribute a very slow Celeron or Pentium, if that would give you some information on how older and slower systems will perform later... as there still are many older units out there.
I would love to see HCF1 protein folding and interaction simulations to help my little boy... someday.
ID: 48380 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
mmonnin

Send message
Joined: 2 Jul 16
Posts: 339
Credit: 8,281,341,558
RAC: 2,803
Level
Tyr
Scientific publications
watwatwatwatwat
Message 48381 - Posted: 15 Dec 2017, 16:15:44 UTC

Quite a few errors on 3.14 so I stopped running QC. Seg faults on AMD and Intel machines.
ID: 48381 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Profile bcavnaugh

Send message
Joined: 8 Nov 13
Posts: 56
Credit: 1,002,640,163
RAC: 0
Level
Met
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwat
Message 48383 - Posted: 15 Dec 2017, 16:43:14 UTC

So I did get some Tasks but it seems that only AMD Processors can run them.
http://www.gpugrid.net/show_host_detail.php?hostid=179948

A nice setting in the Preference for this Tasks would be to allow us to set the number of Cores;
Like on a 32 Core Host you could set 2 Tasks running 16 Cores Each.
Or even 2 Tasks running 8 Cores Each.

Crunching@EVGA The Number One Team in the BOINC Community. Folding@EVGA The Number One Team in the Folding@Home Community.
ID: 48383 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
el_gallo_azul

Send message
Joined: 14 Jun 14
Posts: 9
Credit: 28,094,797
RAC: 0
Level
Val
Scientific publications
watwatwat
Message 48384 - Posted: 16 Dec 2017, 8:20:48 UTC
Last modified: 16 Dec 2017, 8:21:40 UTC

I received ~60 new WUs yesterday, but I didn't see what happened with them. I was surprised when I went back to the computer an hour or so later and they had all disappeared.

I received another batch of ~60 WUs today, and this time I see that they all resulted in "Computation error".

Intel Xeon E5-2680 x 2 (ie. 32 hyperthreading cores).
ID: 48384 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
1 · 2 · 3 · 4 . . . 7 · Next

Message boards : Multicore CPUs : New batch of QC tasks (QMML)

©2026 Universitat Pompeu Fabra