I’ve had to draw a line under the Hotsos paper (not least because I’m supposed to be going out tonight to someone’s birthday celebration meal). So I’ve uploaded the first proper version to both the Symposium website and here.
For those poor souls who have had a look at earlier versions, most of the changes are cosmetic or fixing bugs, but the Multi-user Tests section starting on page 31 until the end of the document is the bit that’s changed the most. This version also has a date on it, so you’ll know you’re looking at the right document!
I need a breather before the presentations …. (although I’m still running some new volume tests through this week but they won’t make it into the paper for a while)
Now you can feel free to criticise 😉
P.S. I think it’s much easier to read if you have a good quality colour print-out than reading it off the screen …
If I find a fatal flaw in the logic, shall I email that privately or bring it up in front of an audience of 500 at the symposium? …
😄
I only printed out the first version yesterday, I was planning to read through it today!
Hope the presentation goes well.
Andy
Well I’m kind of hoping that JL speaking at the same time in the other room should take care of all of the real wizards 😉
Nah, I wouldn’t worry too much – I’ve spotted quite a few flaws myself, as you’ve probably noticed, but the tests still do what they do and show what they show and I don’t regret a thing (gosh, I can feel a song coming on here …)
See you there,
Doug
(I’m writing this in a hurry as I have an appointment with some beer and a thai curry!)
My first simple run through of the benchmark on our production 10gR1 system seemed to suggest that a DOP much beyond 4 – 6 on our system was not having that much extra benefit – in some cases elapsed time was falling.
Obviously I’m not posting results and detailed trace files here and there is far more analysis (for me) to do, but if these results are indicative, they show two things so far:
1. Our assumed DOP of 16 as a sweetspot on our system is too high – we’d be better at somewhere between 4 and 6 and I guess the magic of 2 means 4 would be OK for us. Is this difference between DOP 2 and 4 worthwhile ? (69 v 41s and 140 v 94s for FTS and HJ tests) – seems like a reasonable improvement to justify DOP 4 instead of 2 perhaps.
2. They show that even if you have a large number of CPUs (24 in our case) it won’t necessarily make much / any difference to performance from increasing the DOP to large numbers. In this sense they seem to support what you are saying…that even though you’ve got CPU to burn you’re not necessarily going to gain anything by trying to use it…maybe it would be better getting those CPUs doing something else at the same time instead.
I tried DOP 1 – 12 inclusive and 16,24,32 and 64.
I ran the single full table scan and the 2 table hash join using the rolling.sh script
There will be differences in our environments (of course) and I may be able to make further knowledge available next week after some analysis.
Good luck at the conference mate.
Jeff,
Thanks for that. I would have replied sooner but have been at Parent’s Evening. Which was much more fun than beer and curry, obviously 😉
I assume you were using the PCTFREE 10 version that actually does some work? I’ve been running it this week on the E10K and I’m getting similar results to you. I’ve also heard from someone involved in benchmark at a vendors benchmarking site who got similar results. It’s far from conclusive proof of anything, but interesting nonetheless.
I think the thing is how quickly the benefits diminish so, the sweet spot might be 2, 4 or 6 for example, but increases in DOP above these pretty low numbers don’t give the benefits you’d expect. That’s to say nothing of the additional resources being used.
Thanks loads for giving it a try.
Cheers,
Doug
yes, I believe I was using the PCTFREE 10 version – I’ll confirm on Monday when back in the office.
Yes – I think that’s the main thing coming out of all this…the benefits diminish quickly as you raise the DOP.
just done a bit more analysis of the trace files from my first benchmark run – notable points are:
1. Majority of wait time (99.99% in all cases) when running DOP of more than 1 is for PX Deq: Execute Reply event. Is this hiding an IO wait ?
2. No direct path read or any notable time spent on other PX idle events.
3. My DOP 64 run only achieved DOP of 40 – not sure but probably cos others using some of the available slaves. All other DOP runs got the DOP they requested.
I need to do another run on Monday and try to gain some sar/vm stats so I can see how the OS/machine is working during the runs.
Jeff,
“1. Majority of wait time (99.99% in all cases) when running DOP of more than 1 is for PX Deq: Execute Reply event. Is this hiding an IO wait ?”
I would say so. It’s an idle event saying – “I’m waiting for data to come back”. Is that all you’re getting at all DOPs? If you’re using the PCTFREE 10 example, I’d exepect to see a wide range of other events, particularly at higher DOPs. I’ll mention some details later (should be in bed really)
“3. My DOP 64 run only achieved DOP of 40 – not sure but probably cos others using some of the available slaves. All other DOP runs got the DOP they requested.”
It didn’t get downgraded by parallel_adaptive_multi_user did it? Although maybe the slaves just weren’t available, as you say.
Actually, I don’t know how well this will come out, but here’s a sample of tkprof output fro the ISP4400 with 4 CPUs switched on, DOP 4, Hash Join PCTFREE 10
Elapsed times include waiting on following events:
Event waited on Times Max. Wait Total Waited
That was a rip-roaring success, wasn’t it!
😉
I’ll post it properly later ….
Can’t check the PCTFREE 10 til Monday…the V$PQ_TQSTAT showed I got the DOP requested in all cases except 64 which only got 40.
No other PX waits at all.
It was PCTFREE 90! Doh!
I’ll repeat the tests with PCTFREE 10 later…results to follow.
Sorry!
Jeff said
“Sorry!”
Don’t be.
1) You’re going to the trouble of trying this out too.
2) It confirms what I was thinking (I was getting worried about your lack of wait events)
3) Now that someone else has had the same Doh! moment, it makes me feel less bad about it 😉 (In fairness though, it’s only because you used the same flawed scripts)
Cheers!
OK…as per earlier email to your yahoo…I’ve rerun with PCTFREE 10 and it seemed to indicate that DOP was the cutoff point beyond which increasing DOP gave only marginal benefits.
Again though, I did see some PX Deq: Signal ACK waits but most of the wait time was still the PX Deq: Execute Reply “idle” wait event.
The queries run too quickly (on our box) – max 16s for the full table scan and 90s for the Hash Join – I think these queries should run much longer to see more realistic differences…so I’ll try increasing the volume if I have the space.
I’m going to add 10032/33/53 to the rolling.sh to get more information (overload) for the next run.
I have more stuff to analyse and I’ll get round to that later tonight.
Oh – can you get a lilac hat…Pete Scott needs one to go with his suit! 😉
Hi Jeff,
Yeah those timings seem wildly different to mine, even on the E10K.
I’ll fire through a couple of log files by email so that we have something more detailed to compare.
Starts to get interesting, this, doesn’t it?
😉
Cheers,
Doug
“Oh – can you get a lilac hat…Pete Scott needs one to go with his suit! ;-)”
Damn! I got green! Which will become clearer later …
Good luck today!
OK – last nights run was with 65m rows in the tt1/2 tables.
Timings are longer now with full table scans ranging from 25s to 210s and hash join ranging from 110s to 1000s.
Full scan marginal benefit seems to tail off around DOP 4 and hash join is now at about DOP 5 – 8.
Again, waits are all PX Deq: Execute reply in the main with a fraction of PX Deq: Signal ACK. No other significant waits at all.
Gotta do some paid work now…but more to look at later.