New batch of QC tasks (QMML)

Message boards : Multicore CPUs : New batch of QC tasks (QMML)
Message board moderation

To post messages, you must log in.

Previous · 1 · 2 · 3 · 4 · 5 · 6 · 7 · Next

AuthorMessage
Profile Conan

Send message
Joined: 25 Mar 09
Posts: 25
Credit: 582,385
RAC: 0
Level
Gly
Scientific publications
wat
Message 48445 - Posted: 20 Dec 2017, 9:07:10 UTC - in response to Message 48444.  
Last modified: 20 Dec 2017, 9:09:36 UTC

I finally finished a task so I can post now.

Can someone explain what the QC app shows for Status in the BOINC Manager. I had a app_config.xml loaded to limit the number of cpu cores it was supposed to use to 4. However in the Status column it showed 16C for the number of cores allotted.

Is that normal? Is that just how it describes itself to BOINC or was it really using all 16 cores?

This was my app_confg.xml
<app_config>
<app>
<name>acemdlong</name>
<max_concurrent>1</max_concurrent>
<gpu_versions>
<gpu_usage>1</gpu_usage>
<cpu_usage>1</cpu_usage>
</gpu_versions>
</app>
<app>
<name>acemdshort</name>
<max_concurrent>1</max_concurrent>
<gpu_versions>
<gpu_usage>1.0</gpu_usage>
<cpu_usage>1</cpu_usage>
</gpu_versions>
</app>
<app>
<name>QC</name>
<max_concurrent>1</max_concurrent>
</app>
<app_version>
<app_name>QC</app_name>
<plan_class>mt</plan_class>
<avg_ncpus>4</avg_ncpus> 
<cmdline>--nthreads 4</cmdline>
</app_version>
</app_config>


Does anyone see anything wrong with the app_config?


Try <avg_ncpus>4.000000</avg_ncpus>

The red highlighted line above, it works for me.

Conan
ID: 48445 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Jim1348

Send message
Joined: 28 Jul 12
Posts: 819
Credit: 1,591,285,971
RAC: 0
Level
His
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 48446 - Posted: 20 Dec 2017, 9:08:52 UTC - in response to Message 48444.  
Last modified: 20 Dec 2017, 9:21:33 UTC

Can someone explain what the QC app shows for Status in the BOINC Manager. I had a app_config.xml loaded to limit the number of cpu cores it was supposed to use to 4. However in the Status column it showed 16C for the number of cores allotted.

Is that normal? Is that just how it describes itself to BOINC or was it really using all 16 cores?

The app_config looks the same as mine. Did you reboot in order to activate it?

In some of these multi-core projects, the Status is not updated until the next group of work units comes in after you have set the app_config. But a reboot usually fixes it.
ID: 48446 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Toni
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist

Send message
Joined: 9 Dec 08
Posts: 1006
Credit: 5,068,599
RAC: 0
Level
Ser
Scientific publications
watwatwatwat
Message 48449 - Posted: 20 Dec 2017, 10:46:22 UTC - in response to Message 48446.  

To clarify run times: all the QMML314rst wus are the same length. Even on a single core, they should not take longer than 20h maximum (on a relatively modern PC). The HTTP messages indicate a connectivity problem of course. I hope they cause a failure soon rather than remaining stuck. Re SElinux... I hope it leaves us in peace.
ID: 48449 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Sebastian M. Bobrecki

Send message
Joined: 4 Oct 09
Posts: 6
Credit: 110,801,812
RAC: 0
Level
Cys
Scientific publications
watwatwatwatwatwatwatwat
Message 48454 - Posted: 20 Dec 2017, 13:27:54 UTC
Last modified: 20 Dec 2017, 13:38:38 UTC

After about 10h and reaching 69.568% app started to use only one core. What's worst it stays in that state for another 10h and perf is indicating that it's in OMP spinlock:

83.49% python libiomp5.so [.] __kmp_wait_yield_4
6.76% python libiomp5.so [.] __kmp_eq_4
5.74% python libiomp5.so [.] __kmp_yield
0.66% python [kernel.vmlinux] [k] entry_SYSCALL_64
...

Edit: On second machine it looks similar but after 6h and 78.698% it stays in that state for about 11h now. Perf:

84.60% python libiomp5.so [.] __kmp_wait_yield_4
6.80% python libiomp5.so [.] __kmp_eq_4
5.77% python libiomp5.so [.] __kmp_yield
0.59% python [kernel.vmlinux] [k] entry_SYSCALL_64
0.37% python [kernel.vmlinux] [k] __schedule
...
ID: 48454 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Jim1348

Send message
Joined: 28 Jul 12
Posts: 819
Credit: 1,591,285,971
RAC: 0
Level
His
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 48455 - Posted: 20 Dec 2017, 13:33:18 UTC - in response to Message 48449.  

Even on a single core, they should not take longer than 20h maximum (on a relatively modern PC).

They are not behaving that well at all. I did not have any work units complete yesterday on four machines. That was running two cores each on two i7-3770s and four cores each on an i7-4770 and a Ryzen 1700. These machines are all Ubuntu 16/17, and run 24/7.
http://www.gpugrid.net/results.php?userid=90514

They must loop back at some point, but I will let them run for another couple of days.

By the way, posting is difficult as the website is often unaccessible for a few minutes at a time. Maybe that is related to some of problems some people are having, but I have not looked into it further.

ID: 48455 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Jim1348

Send message
Joined: 28 Jul 12
Posts: 819
Credit: 1,591,285,971
RAC: 0
Level
His
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 48458 - Posted: 20 Dec 2017, 14:54:10 UTC - in response to Message 48455.  

Two have just completed on my Ryzen 1700 (4 cores each). The elapsed time shows as 4 hours 10 minutes, but the CPU time is over two days.
http://www.gpugrid.net/results.php?hostid=452287
ID: 48458 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Keith Myers
Avatar

Send message
Joined: 13 Dec 17
Posts: 1424
Credit: 9,189,946,190
RAC: 0
Level
Tyr
Scientific publications
watwatwatwatwat
Message 48460 - Posted: 20 Dec 2017, 17:10:57 UTC - in response to Message 48446.  

Can someone explain what the QC app shows for Status in the BOINC Manager. I had a app_config.xml loaded to limit the number of cpu cores it was supposed to use to 4. However in the Status column it showed 16C for the number of cores allotted.

Is that normal? Is that just how it describes itself to BOINC or was it really using all 16 cores?

The app_config looks the same as mine. Did you reboot in order to activate it?

In some of these multi-core projects, the Status is not updated until the next group of work units comes in after you have set the app_config. But a reboot usually fixes it.

I reloaded the app_config via the Manager. I was afraid to reboot the machine because I had read earlier in the thread that the tasks would restart and I would lose the processing up to that point. It is normal for BOINC to identify downloaded tasks with the existing cpu/gpu resource usage at time of download. But I could swear I had the app_config in place before I finally snagged my first two tasks. Will wait and see when I can get my next task.
ID: 48460 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Jim1348

Send message
Joined: 28 Jul 12
Posts: 819
Credit: 1,591,285,971
RAC: 0
Level
His
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 48464 - Posted: 20 Dec 2017, 21:02:56 UTC - in response to Message 48460.  

But I could swear I had the app_config in place before I finally snagged my first two tasks. Will wait and see when I can get my next task.

You have to activate the app_config. If you have BoincTasks, there is a way to read all the cc_config and app_config files for any connected machine. (I don't have it in front of me at the moment). Otherwise, a reboot will be necessary. I am having all sorts of problems with the work units, and a reboot is probably not worse than anything else at the moment.
ID: 48464 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
mmonnin

Send message
Joined: 2 Jul 16
Posts: 339
Credit: 8,281,341,558
RAC: 2,803
Level
Tyr
Scientific publications
watwatwatwatwat
Message 48465 - Posted: 20 Dec 2017, 23:11:28 UTC - in response to Message 48464.  

The official BOINC Manager has an option to reread config files as well. I use BOINC Tasks and a new/updated file is picked up without a reboot or client restart.
ID: 48465 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Profile Conan

Send message
Joined: 25 Mar 09
Posts: 25
Credit: 582,385
RAC: 0
Level
Gly
Scientific publications
wat
Message 48467 - Posted: 21 Dec 2017, 3:40:51 UTC

There is also an option in the BOINC code that allows for the number of cpus that you want to use per host.
Each host can have a different setting.

Ask over at Amicable Numbers as they have done that there.
Then you wont need app_config.xml files at all.

Conan
ID: 48467 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Keith Myers
Avatar

Send message
Joined: 13 Dec 17
Posts: 1424
Credit: 9,189,946,190
RAC: 0
Level
Tyr
Scientific publications
watwatwatwatwat
Message 48469 - Posted: 21 Dec 2017, 6:26:09 UTC - in response to Message 48467.  

Setting CPU % in BOINC is system and project wide. Not very good for fine tuning per project. The app_config was specifically introduced for specific project tuning and is the preferred method to control gpu and cpu usage per application.

I have the cpu cores limited in my app_config for both the ACEMD and QC apps. I was wondering why it wasn't picked up after all config files were re-read.
ID: 48469 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Profile Conan

Send message
Joined: 25 Mar 09
Posts: 25
Credit: 582,385
RAC: 0
Level
Gly
Scientific publications
wat
Message 48471 - Posted: 21 Dec 2017, 12:02:09 UTC - in response to Message 48467.  
Last modified: 21 Dec 2017, 12:03:40 UTC

There is also an option in the BOINC code that allows for the number of cpus that you want to use per host.
Each host can have a different setting.

Ask over at Amicable Numbers as they have done that there.
Then you wont need app_config.xml files at all.

Conan


Setting CPU % in BOINC is system and project wide. Not very good for fine tuning per project. The app_config was specifically introduced for specific project tuning and is the preferred method to control gpu and cpu usage per application.

I have the cpu cores limited in my app_config for both the ACEMD and QC apps. I was wondering why it wasn't picked up after all config files were re-read.
Keith Myers


I am not referring to the Boinc Client on your personal computer.
My comments were aimed at the BOINC Server Code and therefore are relevant to fine tuning per project.
The option I am referring to is meant for Multiple Threading, so you can set the number of cores that you want to run a MT work unit on.

Over at Amicable Numbers I have normal for my 4 Core host, 5 cores for my 6 core host and 8 cores for my 16 core host (allowing 2 work units to run at the same time), so that none of the computers have the same setting but could if I wanted them to.

Had the same issue I found here with the 16 core machine and that is why I set it to 8 cores.

App_config.xml works and works well, I was offering an option especially for those of us that are not too good at creating these xml files, and to show that BOINC does have an option in its code to cover this situation.

Conan
ID: 48471 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Jim1348

Send message
Joined: 28 Jul 12
Posts: 819
Credit: 1,591,285,971
RAC: 0
Level
His
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 48472 - Posted: 21 Dec 2017, 12:35:13 UTC - in response to Message 48471.  

I am not referring to the Boinc Client on your personal computer.
My comments were aimed at the BOINC Server Code and therefore are relevant to fine tuning per project.
The option I am referring to is meant for Multiple Threading, so you can set the number of cores that you want to run a MT work unit on.

They have that at LHC too, for the ATLAS project. And they used to do something similar at WCG for the CEP2 project (though that was not mt), in order to limit the high number of writes to the disk drive.

I think it would be very valuable here, since it appears that limiting the number of cores will be needed for many people, and not everyone will be willing to use app_config.xml files.
ID: 48472 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Toni
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist

Send message
Joined: 9 Dec 08
Posts: 1006
Credit: 5,068,599
RAC: 0
Level
Ser
Scientific publications
watwatwatwat
Message 48473 - Posted: 21 Dec 2017, 13:04:51 UTC - in response to Message 48472.  

We'll try to limit the number of cores indeed. It requires server-side changes so may not be soon & may not work.
ID: 48473 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Keith Myers
Avatar

Send message
Joined: 13 Dec 17
Posts: 1424
Credit: 9,189,946,190
RAC: 0
Level
Tyr
Scientific publications
watwatwatwatwat
Message 48476 - Posted: 21 Dec 2017, 19:11:10 UTC - in response to Message 48471.  

I understood what you are saying. The only way that works is if you set up different venues for different projects. I work primarily at SETI and Einstein. The venue mechanism does not work correctly and will likely never be updated. Very low chance that any major rework of the BOINC server code happens in the future with the lack of developers.

I am very comfortable with writing and editing app_info and app_config. Been doing it for a very long while. App_config is the simplest way to tune for individual projects as long as you are using a later version of the Client.

I also run more than one project simultaneously which makes your solution unworkable.
ID: 48476 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
ETQuestor

Send message
Joined: 11 Jul 09
Posts: 27
Credit: 1,000,618,568
RAC: 0
Level
Met
Scientific publications
watwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwatwat
Message 48477 - Posted: 21 Dec 2017, 19:41:07 UTC
Last modified: 21 Dec 2017, 19:44:41 UTC

I just had to restart the BOINC client and the QMML work unit started back from 0% "fraction done" even though it had a checkpoint time of ~130000 and was at ~65%. Boo.

name: c457-TONI_QMML314rst-0-1-RND3080_4
WU name: c457-TONI_QMML314rst-0-1-RND3080
project URL: http://www.gpugrid.net/
received: Thu Dec 21 01:43:45 2017
report deadline: Tue Dec 26 01:43:44 2017
ready to report: no
got server ack: no
final CPU time: 0.000000
state: downloaded
scheduler state: scheduled
exit_status: 0
signal: 0
suspended via GUI: no
active_task_state: EXECUTING
app version num: 314
checkpoint CPU time: 135211.300000
current CPU time: 138876.400000
fraction done: 0.010989
swap size: 1236 MB
working set size: 304 MB
estimated CPU time remaining: 747697.676876
ID: 48477 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Toni
Volunteer moderator
Project administrator
Project developer
Project tester
Project scientist

Send message
Joined: 9 Dec 08
Posts: 1006
Credit: 5,068,599
RAC: 0
Level
Ser
Scientific publications
watwatwatwat
Message 48483 - Posted: 22 Dec 2017, 21:45:12 UTC - in response to Message 48477.  

Can anybody please try if a couple of tasks can run in simultaneously?
ID: 48483 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Keith Myers
Avatar

Send message
Joined: 13 Dec 17
Posts: 1424
Credit: 9,189,946,190
RAC: 0
Level
Tyr
Scientific publications
watwatwatwatwat
Message 48500 - Posted: 24 Dec 2017, 17:15:46 UTC - in response to Message 48483.  

If another one would be made available, I could try. I only ever see one task ready to be snagged. Just got one. Happy to report the change in allowed cores limit was properly applied after I changed 4 to 4.0000.
ID: 48500 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
bormolino

Send message
Joined: 16 May 13
Posts: 41
Credit: 145,731,947
RAC: 0
Level
Cys
Scientific publications
watwatwatwatwatwat
Message 48511 - Posted: 26 Dec 2017, 13:56:46 UTC

Why is there such a big credit difference?



16794335 12932706 25 Dec 2017 | 10:46:18 UTC 26 Dec 2017 | 5:21:03 UTC Fertig und Bestätigt 66,830.38 199,504.30 440.38 Quantum Chemistry v3.14 (mt)

16792051 12932673 24 Dec 2017 | 10:37:41 UTC 25 Dec 2017 | 5:22:21 UTC Fertig und Bestätigt 67,429.09 200,424.90 1,627.77 Quantum Chemistry v3.14 (mt)
ID: 48511 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Keith Myers
Avatar

Send message
Joined: 13 Dec 17
Posts: 1424
Credit: 9,189,946,190
RAC: 0
Level
Tyr
Scientific publications
watwatwatwatwat
Message 48512 - Posted: 27 Dec 2017, 3:04:31 UTC - in response to Message 48483.  

I just grabbed 2 QC tasks and I will attempt to run them simultaneously tomorrow during the SETI outage.
ID: 48512 · Rating: 0 · rate: Rate + / Rate - Report as offensive     Reply Quote
Previous · 1 · 2 · 3 · 4 · 5 · 6 · 7 · Next

Message boards : Multicore CPUs : New batch of QC tasks (QMML)

©2026 Universitat Pompeu Fabra