Message boards :
Multicore CPUs :
New batch of QC tasks (QMML)
Message board moderation
| Author | Message |
|---|---|
|
Send message Joined: 9 Dec 08 Posts: 1006 Credit: 5,068,599 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() |
These are called QMML, and rather experimental (more dependencies). Let's see how they work. |
|
Send message Joined: 9 May 13 Posts: 171 Credit: 4,739,796,466 RAC: 1,182 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
Toni, I have one of these that has been running for over 2 hours and so far looks like it is only using one CPU (thread). It has 4 threads allocated to each task. There are also a number of warnings messages in the stderr.txt that look like this: /var/lib/boinc-client/projects/www.gpugrid.net/miniconda/envs/qmml/lib/python3.6/site-packages/tables/path.py:112: NaturalNameWarning: object name is not a valid Python identifier: '122'; it does not match the pattern ``^[a-zA-Z_][a-zA-Z0-9_]*$``; you will not be able to use natural naming to access this object; using ``getattr()`` will still work, though You should be able to see all of them when the task uploads. Let me know if you need more info. |
|
Send message Joined: 9 Dec 08 Posts: 1006 Credit: 5,068,599 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() |
Thanks, the warnings are expected and harmless. The thread allocation has some bug. I would have expected to use more threads than allocated, not less, but hey. I'll be debugging. Also I hope suspend/resume and the progress bar work (more or less), unlike the old "plain" QC tasks. |
|
Send message Joined: 17 Feb 09 Posts: 91 Credit: 1,603,303,394 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
Subj: Observation of QC CPU WU's Core Utilization Regarding the last go around several weeks ago with the QC CPU WU's that ran successfully on my 8 and 4 core cpu's, the following was observed. 4-core (Phenom II 3 GHz) cpu: core utilization was consistantly near 100% all 4 cores. Work completed in about 40 minutes per WU calender time. 8-core (FX-8350 4 GHz) cpu's (2): 4-cores (alternatively) were utilized at 100% with the remaining cores at much lesser utilization. Work completed with about 20 minutes real time per WU. With the most recent QC WU's (12/13/2017), one FX-8350 errors out consistantly and will not run the WU's (maybe it requires software or drivers not currently installed) and the other FX-8350 runs them but at a substanially less core utilization rate than the prior WU batches. So far, the first WU in progress is currently less than 40% complete with about 2 hours calender time invested. Not sure if these recent WU's are of the same length as the earlier ones however. I have uploaded photos of ksysguard graphic cpu utilization that can be observed if interested with addresses below. Basically, as I am sure that this is a work in progress and bugs need to be resolved but I would conclude so far that these WU's process efficiently on a 4-core system but that all 8-cores should be fully utilized to make it worth while sacrificing 8-cores to a single WU when 8 WU's from other projects use all 8 at 100% thereby being much more efficient. Haven't tried the suspend/resume yet but will when the opprotunity is available. The previous poster appears correct re thread utilization being one issue. Screenshots: http://members.toast.net/obc/computing/grid_computing/images/QC_4-core_cpu.png http://members.toast.net/obc/computing/grid_computing/images/QC_8-core_cpu.png http://members.toast.net/obc/computing/grid_computing/images/QC-12-13-2017.png (Sorry, crude way to present photos but wanted to get them up for anyone interested) |
|
Send message Joined: 2 Jul 16 Posts: 339 Credit: 8,281,341,558 RAC: 2,803 Level ![]() Scientific publications ![]() ![]() ![]() ![]()
|
Mine is also just a bit less than 1 thread with 7 available for CPU usage. Good thing I have another client available to keep the CPU busy. If the progress bar is correct it will take about 9.5 hours on 1950x at 3.75 GHz. |
|
Send message Joined: 9 Dec 08 Posts: 1006 Credit: 5,068,599 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() |
The next batch (QMML313a) should respect the number of threads requested by your client. |
ConanSend message Joined: 25 Mar 09 Posts: 25 Credit: 582,385 RAC: 0 Level ![]() Scientific publications
|
The Multiple threaded work units that were sent out last month worked fine for me with no issues. These new ones however are all failing (so far on two different 64bit Linux machines) with this error ERROR conda.core.link:_execute_actions(337): An error occurred while installing package 'psi4::gcc-5-5.2.0-1'. LinkError: post-link script failed for package psi4::gcc-5-5.2.0-1 running your command again with `-v` will provide additional information location of failed script: /home/Conan/BOINC/projects/www.gpugrid.net/miniconda/envs/qmml/bin/.gcc-5-post-link.sh When checking this path I found that there is nothing in the /envs/ folder, which is probably where the job is failing. Conan |
|
Send message Joined: 2 Jul 16 Posts: 339 Credit: 8,281,341,558 RAC: 2,803 Level ![]() Scientific publications ![]() ![]() ![]() ![]()
|
Hmm 1st one completed for me and the 2nd one is at around 86%. |
|
Send message Joined: 9 May 13 Posts: 171 Credit: 4,739,796,466 RAC: 1,182 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
Toni said, The next batch (QMML313a) should respect the number of threads requested by your client. It looks like this one does respect the number of threads requested by my client. My app_config specifies 4 threads and it looks to be using 4 threads. Let me know if you need more info. |
|
Send message Joined: 28 Jul 12 Posts: 819 Credit: 1,591,285,971 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
I am not able to get QC on my Ryzen 1700 machine running Ubuntu 17.10. I just get a "No tasks are available for Quantum Chemistry" message when I request them. However, I am able to get QC on my i7 3770 machine running Ubuntu 16.04 (both machines have BOINC 7.8.3). Both machines are set to the same profile (work), so they should be treated identically. But I see that some people with AMD machines get work. Is this a bug or a feature? |
|
Send message Joined: 9 Dec 08 Posts: 1006 Credit: 5,068,599 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() |
The reason why otherwise similar machines get/do not get work completely baffles me. I don't think it's related to the maker of the CPU. Perhaps with the history of tasks/host reliability or somesuch. With this respects BOINC is of no help. |
bcavnaughSend message Joined: 8 Nov 13 Posts: 56 Credit: 1,002,640,163 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
These are called QMML, and rather experimental (more dependencies). Let's see how they work. I would really like to get some tasks but seems they are not being given out ATM http://www.gpugrid.net/show_host_detail.php?hostid=457056 Been Trying for awhile now. Intel(R) Core(TM) i7-3970X Crunching@EVGA The Number One Team in the BOINC Community. Folding@EVGA The Number One Team in the Folding@Home Community. |
|
Send message Joined: 28 Jul 12 Posts: 819 Credit: 1,591,285,971 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
The reason why otherwise similar machines get/do not get work completely baffles me. I have seen many instances of it myself over the years, but had hoped the latest BOINC clients were past that. Unfortunately not. |
|
Send message Joined: 2 Jul 16 Posts: 339 Credit: 8,281,341,558 RAC: 2,803 Level ![]() Scientific publications ![]() ![]() ![]() ![]()
|
My 1950x also on 17.10 is getting tasks. Were the old 3.13 tasks producing bad data as my processing time was just wasted by the server canceling them in the middle of processing them? Thats the absolute worst thing a project admin can do. Cancel ones not started but don't ever take a crap on donated resources. Its still not working right. In 22.5min of the task running it has used 1:37min of CPU time when the task is limited to 3 cores. That's over 4 cores of CPU usage. And they just had a computation error. At least 3.13 worked. http://www.gpugrid.net/result.php?resultid=16767178 Prob a good thing tasks are being sent to some AMD CPUs. Damn seg fault. |
|
Send message Joined: 23 Dec 09 Posts: 189 Credit: 4,813,881,008 RAC: 149 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
Most of the Quantum Chemistry v3.14 (mt) fail on my AMD 1700x. v3.13 (mt) worked more or less. As an example: http://www.gpugrid.net/result.php?resultid=16767309 I use an app_config to limit the use to 4 cores for each work unit (WU) and runs 3 WUs in parallel. Two cores are reserved for the GPU. I had to change the configuration to accept only GPU Work Unites as the computer crashed twice today. Hope this helps. |
|
Send message Joined: 9 Dec 08 Posts: 1006 Credit: 5,068,599 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() |
I understand that seing cancelled WUs is not nice, but it saves future crunching and network bandwidth (both server and client) that would otherwise be lost. Also, I thought that the function we use only cancelled UNSENT or un started wus. |
|
Send message Joined: 20 Apr 15 Posts: 285 Credit: 1,102,216,607 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]()
|
Hey friends, in case you need some more machines for testing, I can set up another one that is Linux based. As I have seen in the below conversation, there might be some issues with the CPU type. So .. I have both brands available for testing. Which one would help you most, Intel or AMD? If you want me to, I could even let you choose the generation. From older Sandy/Ivy Bridge to new Skylake to Ryzen. Lust let me know and I will make one available on short notice. Edit: I can even contribute a very slow Celeron or Pentium, if that would give you some information on how older and slower systems will perform later... as there still are many older units out there. I would love to see HCF1 protein folding and interaction simulations to help my little boy... someday. |
|
Send message Joined: 2 Jul 16 Posts: 339 Credit: 8,281,341,558 RAC: 2,803 Level ![]() Scientific publications ![]() ![]() ![]() ![]()
|
Quite a few errors on 3.14 so I stopped running QC. Seg faults on AMD and Intel machines. |
bcavnaughSend message Joined: 8 Nov 13 Posts: 56 Credit: 1,002,640,163 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
So I did get some Tasks but it seems that only AMD Processors can run them. http://www.gpugrid.net/show_host_detail.php?hostid=179948 A nice setting in the Preference for this Tasks would be to allow us to set the number of Cores; Like on a 32 Core Host you could set 2 Tasks running 16 Cores Each. Or even 2 Tasks running 8 Cores Each. Crunching@EVGA The Number One Team in the BOINC Community. Folding@EVGA The Number One Team in the Folding@Home Community. |
|
Send message Joined: 14 Jun 14 Posts: 9 Credit: 28,094,797 RAC: 0 Level ![]() Scientific publications ![]() ![]()
|
I received ~60 new WUs yesterday, but I didn't see what happened with them. I was surprised when I went back to the computer an hour or so later and they had all disappeared. I received another batch of ~60 WUs today, and this time I see that they all resulted in "Computation error". Intel Xeon E5-2680 x 2 (ie. 32 hyperthreading cores). |
©2026 Universitat Pompeu Fabra