Message boards :
Multicore CPUs :
New batch of QC tasks (QMML)
Message board moderation
Previous · 1 · 2 · 3 · 4 · 5 · 6 · 7 · Next
| Author | Message |
|---|---|
|
Send message Joined: 17 Feb 09 Posts: 91 Credit: 1,603,303,394 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
These two WU's ran concurrently on one of my FX8350 using 4 cores each. They were started within about twenty minutes of each other and finished about 4 minutes apart. 16795026 12932509 426610 25 Dec 2017 | 15:47:53 UTC 27 Dec 2017 | 12:23:12 UTC Completed and validated 53,395.91 188,995.80 3,034.10 Quantum Chemistry v3.14 (mt) 16795025 12932492 426610 25 Dec 2017 | 15:47:53 UTC 27 Dec 2017 | 12:19:03 UTC Completed and validated 53,957.92 193,510.10 3,066.04 Quantum Chemistry v3.14 (mt) |
|
Send message Joined: 13 Dec 17 Posts: 1424 Credit: 9,189,946,190 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]()
|
I just grabbed 2 QC tasks and I will attempt to run them simultaneously tomorrow during the SETI outage. I just finished these two QC tasks run concurrently. 16798843 12932712 456812 27 Dec 2017 | 3:00:54 UTC 28 Dec 2017 | 7:57:47 UTC Completed and validated 26,162.47 103,683.40 4,277.76 Quantum Chemistry v3.14 (mt) 16798838 12932759 456812 27 Dec 2017 | 2:58:51 UTC 28 Dec 2017 | 7:57:47 UTC Completed and validated 26,201.39 103,805.90 4,284.12 Quantum Chemistry v3.14 (mt) I started them within a minute of each other using 4 cores each. I also had 3 Einstein GPU tasks running concurrently with them. System is a AMD Ryzen 1800X 16 core CPU and three Nvidia GTX 970's. Didn't appear to have any problems. Tasks ran right through with about 70% CPU utilization. |
|
Send message Joined: 9 Dec 08 Posts: 1006 Credit: 5,068,599 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() |
Thanks @keith, thanks @starbase! |
|
Send message Joined: 13 Dec 17 Posts: 1424 Credit: 9,189,946,190 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]()
|
Toni, well there's your answer if you discount the low population sample. I doubt that two successful runs were due to the brand of cpu, but could be wrong. As long as you are using a Client later than 7.0.40 you can use an app_config.xml file to tune the number of cores you allow the task to run on. |
|
Send message Joined: 28 Jul 12 Posts: 819 Credit: 1,591,285,971 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
My Ryzen 1700 machine has certainly done better than my two i7-3770 PCs (all machines on Ubuntu, run 24/7 and otherwise set up the same): Ryzen 1700: http://www.gpugrid.net/results.php?hostid=452287 i7-3770: http://www.gpugrid.net/results.php?hostid=433866 http://www.gpugrid.net/results.php?hostid=448995 EDIT: My i7-4770 PC also tended to hang or otherwise fail. http://www.gpugrid.net/results.php?hostid=357332 I have often gotten hung work units on the Intel machines, but seldom or never on the AMD. And looking around at the other users who fail the work units, they seem to be predominantly Intel, while the ones that succeed seem to be AMD, though I have not done a count myself. Presumably Toni can get those figures. |
|
Send message Joined: 13 Dec 17 Posts: 1424 Credit: 9,189,946,190 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]()
|
Thanks for the post. Interesting. I could discount the fact that the Ryzen's have actual 8 physical cores so on paper a good head start over the Intel 4 core cpu's. But the FX-9350 earlier in the thread had a good result too with only 4 physical cores too and much more handicapped FFT registers in its modules compared to Ryzen and Intel. Need a lot more samples to definitively clarify I think. |
|
Send message Joined: 28 Jul 12 Posts: 819 Credit: 1,591,285,971 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
Need a lot more samples to definitively clarify I think. Yes. I am not looking at the output, but really only the error rate. All the machines are now running two cores per work unit, and only one work unit per machine, though earlier I had been running four cores on the AMD machine. And both the Intel and AMD cores have about the same speed, so the output should be comparable anyway now. I have changed one of the i7-3770 machines (GTX-1070-PC) from Ubuntu 17.10 (and BOINC 7.8.3) back to Ubuntu 16.04 (and BOINC 7.6.31). I doubt that it will make much difference, but I will let it run for a couple of weeks. If I continue to get more errors on Intel, I think I will go with just the Ryzen PC. |
|
Send message Joined: 13 Dec 17 Posts: 1424 Credit: 9,189,946,190 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]()
|
Good experiment. I was just thinking that spreading the compute load over 4 cores is less hard compared to 100% workload over 4 cores on Intel. If the test on 2 and one cores is equally stable, just slower, it might suggest something bothersome on Intel architecture. The different OS platform could have a big effect too. |
|
Send message Joined: 17 Feb 09 Posts: 91 Credit: 1,603,303,394 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
It appears one cannot start two or more of these type of WU's simultaneously. One of the two errored as shown below. 16803962 12952984 426610 29 Dec 2017 | 16:16:51 UTC 29 Dec 2017 | 17:15:51 UTC Completed and validated 916.51 6,109.03 35.20 Quantum Chemistry v3.14 (mt) 16803949 12952971 426610 29 Dec 2017 | 16:16:51 UTC 29 Dec 2017 | 17:26:30 UTC Error while computing 5.04 0.00 --- Quantum Chemistry v3.14 (mt) So I started two more but with a five second delay and they both are now happily processing together on one of my FX8350's 4 cores each. With that in mind, the boinc client may not be relied upon to run these cpu jobs unattended using a split cpu core configuration to allow multiple WU processing lest two or more were to start simultaneously causing a possible failure on at least one WU. Edit: Will try a simultaneous start with my last two WU's to see if this is repeatable. Edit2: Yep, happened again. Will copy the pair to this post when the one in progress finishes. |
|
Send message Joined: 17 Feb 09 Posts: 91 Credit: 1,603,303,394 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
These are the last two, one of which errored: 16804044 12953066 426610 29 Dec 2017 | 17:28:46 UTC 29 Dec 2017 | 21:02:26 UTC Error while computing 6.06 0.00 --- Quantum Chemistry v3.14 (mt) 16804026 12953048 426610 29 Dec 2017 | 17:28:46 UTC 29 Dec 2017 | 21:24:44 UTC Completed and validated 1,393.35 5,255.89 60.62 Quantum Chemistry v3.14 (mt) |
|
Send message Joined: 28 Jul 12 Posts: 819 Credit: 1,591,285,971 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
It appears one cannot start two or more of these type of WU's simultaneously. That is curious. Thanks for the report. The new DOMINIKs run fine on my two i7-3770s thus far, probably since they are shorter than the TONIs and don't get to the point of hanging up. And the Ryzen 1700 continues to do well. But there must be some selection process going on at the server, since it is the getting only the TONIs. They are all reissues now, but it has handled them all thus far, even a _8. That is a good idea, since it makes optimum use of each CPU type. If things continue this way, I will just let all the machines run. |
|
Send message Joined: 9 Dec 08 Posts: 1006 Credit: 5,068,599 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() |
Looks like that multiple WUs together is ok, but starting exactly at the same time is not. I hope it is a relatively rare occurrence. In principle I could put a locking mechanism, but I am not enthusiastic because that would be inviting more failure modes (e.g. stale locks) to solve a relatively rare case. |
|
Send message Joined: 2 Jul 16 Posts: 339 Credit: 8,281,341,558 RAC: 2,803 Level ![]() Scientific publications ![]() ![]() ![]() ![]()
|
Looks like that multiple WUs together is ok, but starting exactly at the same time is not. I hope it is a relatively rare occurrence. In principle I could put a locking mechanism, but I am not enthusiastic because that would be inviting more failure modes (e.g. stale locks) to solve a relatively rare case. So every new member crunching these will get an error on their 1st task. Genius. It absolutely should be fixed. |
|
Send message Joined: 28 Jul 12 Posts: 819 Credit: 1,591,285,971 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
Looks like that multiple WUs together is ok, but starting exactly at the same time is not. I hope it is a relatively rare occurrence. In principle I could put a locking mechanism, but I am not enthusiastic because that would be inviting more failure modes (e.g. stale locks) to solve a relatively rare case. I got two errors at first also, but none since and I did not look at the reason. But it is over and done with, and not a problem. I think if you look hard at the logic of lock mechanisms, it is a logical impossibility to fix simultaneity problems. That is, any delay you put it will match some other starting situation, and result in an error also. You can try, but I don't think it is worth the effort either. EDIT: I would look to see if it happens again with the next batch. If so, then I would investigate something, whatever it is. But the errors were very short for me, and no real time lost. |
|
Send message Joined: 17 Feb 09 Posts: 91 Credit: 1,603,303,394 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
I would speculate that the probability of two or more WU's starting at exactly the same time in an unattended environment would be low however, those running multiple projects competing for the same GPU/CPU resources are usually switched at a specified time interval to give each project process time. With that in mind, it is possible for two QC jobs to finish and the boinc client nearing the end of a switch app interval, change to the other project and then when it came time to switch back to run the cpu gpugrid WU's with a queue full of QC jobs start two or more simultaneously. The only way I can see avoiding the possibility of such an error completely would be to keep all cores dedicated to one QC job on unattended machines (headless crunchers in my case). When the next batch is available, I will cut my switch app time down to say 5 minutes and see if I can get a feel for how the client handles things and the possibility of simultaneous starts with the QC jobs. The two jobs that failed due to simultaneous starts both ended with exit code 195 EXIT_CHILD_FAILED. Not sure if that means the app failed to spawn a child thread or close one for one of the two WU's. |
|
Send message Joined: 28 Jul 12 Posts: 819 Credit: 1,591,285,971 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
You are right that BOINC is not really a random environment, and if it happens in a more-or-less predictable manner, it should be possible to prevent it. We will see how often that is. |
|
Send message Joined: 17 Feb 09 Posts: 91 Credit: 1,603,303,394 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
You are right that BOINC is not really a random environment, and if it happens in a more-or-less predictable manner, it should be possible to prevent it. We will see how often that is. Agreed, predicable is the key. I really need to find out how the projects I crunch for work unattended because all my little headless crunchers are in process of being converted to diskless/headless cluster nodes with one of the FX system being the master. This should all prove interesting as this is the first project I've worked with that uses multicores for a single WU. |
|
Send message Joined: 2 Jul 16 Posts: 339 Credit: 8,281,341,558 RAC: 2,803 Level ![]() Scientific publications ![]() ![]() ![]() ![]()
|
When there is work or the GPUGrid servers sent out work they come in batches. Two tasks end up starting at once since they have short deadlines. I also said new members so again multiple tasks starting at once. Not everyone runs the same project all the time so that tasks have a chance to get off sequence. |
|
Send message Joined: 28 Jul 12 Posts: 819 Credit: 1,591,285,971 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
I have changed one of the i7-3770 machines (GTX-1070-PC) from Ubuntu 17.10 (and BOINC 7.8.3) back to Ubuntu 16.04 (and BOINC 7.6.31). I just completed two TONI work units, one on each of my i7-3770 PCs (2 cores per work unit): http://www.gpugrid.net/workunit.php?wuid=12932866 http://www.gpugrid.net/workunit.php?wuid=12932333 They had each errored out on other PCs, and I don't know why they worked on mine. But I do know that they each got stuck at 78.698% until I rebooted, and then they completed normally. However, the total Run time shown does not include the time they were stuck, which was about two hours in each case. This is no way to get work done; I can't be rebooting for each work unit. So I will have to just stop on the i7-3770 machines and continue only with the Ryzen 1700, which continues to work fine. Note that the new DOMINIK work units are no problem - if they could send only those to the Intel machines, I think the problem would be solved. |
|
Send message Joined: 13 Dec 17 Posts: 1424 Credit: 9,189,946,190 RAC: 0 Level ![]() Scientific publications ![]() ![]() ![]() ![]()
|
I just ran a DOMINIK QC task and it ran very fast. Task 16815024 Anybody else run one of these yet? I see that there are a TON of them available. Server Status |
©2026 Universitat Pompeu Fabra