PC Magazine: Apple Power Mac G5 Quad delivers ‘performance numbers we’ve never seen before’

“Apple professionals expect a speed bump with each successive Power Mac model that comes out, and the 2.5-GHz Power Mac G5 Quad delivers a doozy of one. With the first Intel-powered consumer Macs rumored to be right around the corner (they may be announced at the MacWorld Expo in early January), it is interesting that Apple would release the G5 Quad as the probable last hurrah for the PowerPC processor—but at least they’re going out on top,” Joel Santo Domingo writes for PC Magazine.

“Is it worth upgrading to the new G5 Quad? The answer is a resounding yes for those who need (and can justify) the power and expense. The dual G5 cores in each of the two CPUs push the G5 Quad to performance numbers we’ve never seen before. With four true cores working on the same task, the G5 Quad powered through our new Adobe Photoshop CS2 tests at a speedy 57 seconds. The previous Power Mac Dual (2.7 GHz) took 1 minute 14 seconds to do the same ten tasks (30 percent longer), and the Dell XPS 600 took 1:03 (a still-significant 11 percent difference). Though 11 percent doesn’t seem like much, it can really add up over the course of a day or week, especially on time-sensitive projects,” Santo Domingo writes.

“The Quad G5 got the highest score we’ve ever seen on the CPU-stressing CineBench rendering test: 1,104. The Pentium EE840 overclocked to 3.6 GHz recently got a 667, and an Athlon 64 4800+ overclocked to 2.7 GHz scored 775,” Santo Domingo writes. “We recommend that professional businesses such as design and engineering firms continue to buy PowerPC-based Power Macs. Intel-native and universal binary software (software that contains both PowerPC and Intel optimized code) are likely to lag behind the introduction of the Intel Macs by several months to a year. Since non-Intel-optimized programs are likely to remain current for several years after the introduction of Intel Macs in 2006 and 2007, it behooves current Mac houses to buy the current PowerPC Macs as their last pre-Intel upgrade.”

Full review (4.5 out of 5 stars) here.

Advertisement: Power Mac G5. Dual-core PowerPC processors with PCI Express. From $1999. Free shipping.

Related MacDailyNews articles:
Benchmarks show Power Mac G5 Quad 2.5GHz a magnificently powerful performance demon – December 19, 2005
Computerworld’s advice on Apple’s new Power Mac Quad G5: Place your orders now! – November 16, 2005
Apple Quad 2.5GHz Power Mac G5 vs. previous generation dual 2.5GHz Power Mac G5 – November 14, 2005
InfoWorld: Nothing can compare to Apple’s new Power Mac G5 Quad – true workstation at desktop price – October 24, 2005
NVIDIA brings workstation graphics to Apple Power Mac G5 – October 24, 2005
Apple’s new Power Mac G5 Quad supercharges rendering – October 22, 2005
AnandTech: Apple new Power Mac G5’s biggest improvement is the move to PCI Express – October 21, 2005
Photos of new dual core Apple Power Mac G5 interior, ports, and more – October 19, 2005
Apple introduces Power Mac G5 Quad and Power Mac G5 Dual – October 19, 2005

61 Comments

  1. Don’t AMD Opterons smoke the Quad G5 on CineBench? Or were they not counting those because they are basically server chips?

    MDN MW: Doubt, as in “I doubt the Quad G5 could touch a Quad Opteron”

  2. Purchased mine earlier this week. Got the upgraded video card, 4GB of ECC RAM, etc. – to the tune of $5,800. My present computer is a Dual G4 1GHz. It’s been really good to me, but I can only imagine how much faster LightWave and After Effects will be.

  3. The only problem with the article… and I quote,

    “Apple professionals expect a speed bump with each successive Power Mac model that comes out, and the 2.5-GHz Power Mac G5 Quad ($9,522 direct, $7,023 without monitor) delivers a doozy of one.”

    That’s the first paragraph… it only coasts 3300. Further down they mention that “the basic configuration of the 2.5-GHz model comes in at a more reasonable $3,299 (without monitor);” ……. . .. .. .

    Um… that’s the topoftheline model… not the basic… and it isn’t 7k… my company wouldn’t have bought 2 if that was the case.

    Am I just reading something wrong?

  4. Oh-my-God … I just have to read that article again:

    “The Quad G5 got the highest score we’ve ever seen on the CPU-stressing CineBench rendering test: 1,104. ” width=”19″ height=”19″ alt=”big surprise” style=”border:0;” />
    The Pentium EE840 overclocked to 3.6 GHz recently got a 667, and an Athlon 64 4800+ overclocked to 2.7 GHz scored 775. CineBench is a multithreaded app, so the more cores or threads your system can handle, the more efficiently your workload gets done.”

    (shaking head in disbelief)

    Note for clarity: Both the EE840 and the 4800+, besides being 200-1100Mhz faster, are dual core CPUs as well – the best that either Intel or AMD currently offer in a desktop that isn’t running server level stuff.

    Look, I know I’m gonna get flamed, but after results like these just keep coming up I can’t help it … Is Intel and Viiv and DRM’d Boobtube 2.0 really, I mean REALLY, worth it?

    ” width=”19″ height=”19″ alt=”cool hmm” style=”border:0;” /> I just don’t think so.

  5. …and the next generation Power Macs with quad Intel chips will again smoke the competition. Apple is quite right moving to Intel. No way would SJ have gone for the move if he was going to go backwards…

  6. Quoting The Mac God
    [“Apple professionals expect a speed bump with each successive Power Mac model that comes out, and the 2.5-GHz Power Mac G5 Quad ($9,522 direct, $7,023 without monitor) delivers a doozy of one.”

    That’s the first paragraph… it only coasts 3300. Further down they mention that “the basic configuration of the 2.5-GHz model comes in at a more reasonable $3,299 (without monitor);]

    The figures they are using are for a setup with 4 gigs of ram, 1 terabyte of harddrive storage and the NVIDIA 512MB video card. Obviously not everyone will order their setup this way but it’s the one they tested.

    Also it should be noted that they chose the 8x512mb ECC RAM. If they did what most smart mac buyers do they would have boosted their ram from crucial, who could get them the extra ram for a total cost of 6442, or nearly $600 less than Apple’s price. That is using 8×512 chips, which is possible but personally i would use 4x1gig chips for upgradability later expecially since right now it would only bring the price to 6445, 3 dollars more.

    Buying memory is a tricky thing, and the trickery can change everytime crucial updates it’s page. I think price comparisons should never include ram prices. Instead they should, as should performance comparisons, note the amount of choice the buyer has ECC, speed, module sizes, and price, making it impossible to say exactly what the cost will be. One thing they should say is that buying RAM direct will save them a lot of money if they are buying large amounts.

  7. >I mean REALLY, worth it?

    The problem is with the portables — a significant market; and intel’s roadmap offered Apple a better opportunity. Sad to say, but my dual 2.5 will have to do until all the kinks are ironed out of Leopard and the new architecture. Till then, I’ve got nothing to complain about. Now as for the media center and drm … won’t be long before we know.

    Signing off for a holiday with Angus and Lazio.

  8. >>Look, I know I’m gonna get flamed, but after results like these just keep coming up I can’t help it … Is Intel and Viiv and DRM’d Boobtube 2.0 really, I mean REALLY, worth it?

    Clearly yes. The Quad G5 (four cores) was marginally faster than a dual-core Intel (two cores). Do the math for a four-core Intel.

  9. Secondly, the benchmarks were again graphics dominated, and the Quad G5 contained the top-of-the-line nVidia Quadro graphics card. Did the other two machines? (I doubt it). So how much of this was a test of the processors and how much of the nVidia Quadro 4500? I note that they could only get 23 fps out of Doom 3, which is way behind current state-of-the-art gaming PC’s. Gamings not everything – but they do represent some of the most stressing all-round applications for a computer.

  10. Better frame that article. We’ll never see it again when we move into the same ‘hood with Dell & Gateway.

    Meanwhile, AMD will continue to spank Intel in benchmarks, while we’re served up some fresh “performance per watt” kool-aid.

    Pfft.

  11. Yeah, it might be fast… but will it eliminate pop-ups and pop-unders on MDN?

    Haven’t seen a popup or popunder in a loong time … nor ads that annoyed me once too often. It’s a cinch. Control click on the ad and filter anything from the source site using a wildcard after the basic url.

    Want to retrive an image, javascript, flash movie from a site? Boom — it’s downloading.

    Want to see that image you saw yesterday but dunno where? Just browse what’s in your cache.

    Don’t want to try another browser because you’re happy with what you’ve got? Fine … live with the poppers. Your choice.

  12. “Look, I know I’m gonna get flamed, but after results like these just keep coming up I can’t help it … Is Intel and Viiv and DRM’d Boobtube 2.0 really, I mean REALLY, worth it”

    I won’t flame you because well what does that do? I’m quite in awe of the specs myself. But the thing is, you are missing the real point of the Intel switch. I’ve posted my theory on this board before so I’ll give you the cliffs version:

    The biggest objection that switchers have to moving to Macs is software availability. There is a real fear there. Switching to Intel removes that fear as it becomes easier for vendors to make mac versions, makes it a better platform to run VPC on, and potentially will either dual-boot, or allow for windows software to run within a window without installing windows (chuckles)

    Therein lies your answer. It has less to do with the DRM argument although I’m sure that will be part of future releases.

  13. Troll: Ick. People still use iCab? Everyone I know gave it up for a lost cause over a year ago. Let them take it out of beta, then I’ll care.

    Odyssey: Clearly the G5 still has some kick it, which is why it will be the last chip to be replaced. Meanwhile, the Yonah looks to completely validate Jobs’ decision when it’s introduced into iBooks very soon.

  14. Hammer:
    Here’s your first flame.

    “The biggest objection that switchers have to moving to Macs is software availability.”

    Blow it out your arse. Every single app my company uses is available on all almost all 3 platforms. Windows, MacOs and Linux (still lagging). If it isn’t it’s not a true “productivity app.” And if you can’t find a similar replacement it’s most likely an app that has seen its day. That argument is sooo lame and soooooo 90’s.

  15. “Look, I know I’m gonna get flamed, but after results like these just keep coming up I can’t help it … Is Intel and Viiv and DRM’d Boobtube 2.0 really, I mean REALLY, worth it?”

    Definitely not worth switching to Intel! One thing is for sure the switch ISN’T about performance if that were the case Apple would be sticking with PowerPC.

    X86 cpu’s also use more watts + heat by design… http://www.osnews.com/story.php?news_id=3997&page=4

  16. “Clearly the G5 still has some kick it, which is why it will be the last chip to be replaced. Meanwhile, the Yonah looks to completely validate Jobs’ decision when it’s introduced into iBooks very soon.”

    Not really Yonah is being compared to 90nm chips. Compare them to future 65nm chips and Intel will lag as usual.

  17. From: Rammer
    Dec 23, 05 – 09:20 am
    max: “Every single app my company uses is available on all almost all 3 platforms.”
    Cool your company uses Internet Explorer, Microsoft Word and Firefox, too!”

    The whole software issue is totally valid. Now, let’s think about the average consumer who is totally clueless about computers. My sister, for example, won’t switch to a mac because of two main reasons.

    1. Her kids use windows in school and she wants to use what they use there (it’s totally bogus but she believes this regardless of any arguement I make.)

    2. When she goes to Target or Wal-mart she can pick up any software title and take it home. She rarely would pick up any software at these places but it makes her feel secure that all of those options are there.

    I have found that those two excuses are the biggest hurdles for macs to overcome. The “We already use windows and all software is made for windows.” attitude is going to be tough to overcome.

    Now, if we can get past the feelings of “software security choice” issue then the mac will most likely get somewhere with a lot of people. Most people like to be able to go to a store, see a software title, and not have to worry if it is going to work on their system or not.

    You can try and sell that mentality on everything else but it won’t work. She has so much adware and spyware on her computer that it is almost unable to be used but she will buy windows again because she can pick up “Hello Kitty does Math” at Target.

  18. >>Secondly, the benchmarks were again graphics dominated, and the Quad G5 contained the top-of-the-line nVidia Quadro graphics card. Did the other two machines? (I doubt it). So how much of this was a test of the processors and how much of the nVidia Quadro 4500? I note that they could only get 23 fps out of Doom 3, which is way behind current state-of-the-art gaming PC’s. Gamings not everything – but they do represent some of the most stressing all-round applications for a computer.

    Sounds like a contradiction.

    Doom 3 is not a valid test for Mac performance. For you psyche majors, the mac version just ain’t the same.

  19. Odyssey67, Macintel, Reality Check and others:

    The move to Intel is Entirely aimed at portables (I’ve posted this before). You want to know what is really impressive?

    Look at these specs for NEC’s next Yonah-based notebook:

    “…with 512MB of main memory and a 100GB hard disk drive. It will have a 14.1-inch LCD, DVD Super Multi drive (DVD-R/+R, DVD-RAM, DVD-RW/+RW), 802.11a/b/g Wi-Fi, and Bluetooth. The machine will weigh about 4.4 pounds and the battery will provide enough power to last about four hours.”

    Compare that to the current 12″ powerbook, weighing in at 4.6 pounds < iBook 12″ 4.9 pounds < 15″ powerbook 5.6 pounds < 14″ iBook 5.9 pounds < 17″ powerbook 6.9 pounds. (All figures with default specs).

    Now if Apple was able to produce the slickest, lightest notebooks before, can you imagine how light the new generation is going to be? Can you imagine having an Apple notebook run on battery power for 5 hours?

    Although the PowerMacs are really powerful, I think I belong to the growing majority of computer consumers that are waiting for the next generation of Apple notebooks.

    Have a great holiday everyone, and a happy New Year!

    Peace…

  20. Went to CNET Shopper to compare prices on that Dell that they used in the comparison to the Quad. First feedback comment from users was this:

    Just got my delivery. 4 YEARS COMPLETE CARE $500. First Thing out of the Box. Tried installing ADOBE 6.0 Full Version.
    Froze for 10 Minutes Finally rebooted. Came about Physical Dump memory. rebooted again. Same message. Got pissed off and called the tech support. THREE HOURS ON THE PHONE. both cordless phones dead. So reinstalled the windows from the cd i got. installation message. Two partitions. first for dell utilities and 2nd for operating system…. installed their. finished installing. 4 hours down the drain so far. went to sleep got up and called tech support. no application cds. some body answered after 2 and half hours. said will send cd’s in mail. will get them in 3 business days. today is the 9 business day no cd’s. told him it keeps crashing…. said let’s install from the image oka…. THEY NEVER CREATED A BACKUP OF MY IMAGE. HOW STUPID ARE THEY.
    PAID $4800 DOLLARS FOR THIS PIECE OF **** TOLD THEM TO COME AND PICK IT UP OR IT’S OUT OF THE WINDOW IN 3 DAYS. IT’S BEEN 9 DAYS NO BODY CAME. PUT A STOP PAYMENT ON THE Credit cards used to buy this. TOLD THEM NEVER BOUGHT IT. Canceled the payment.

    LOL! And they guy paid more for his Dull than I did for my Quad, and it hasn’t had so much as a hiccup.

    MW = yes, as in “Yes, I’m very happy with my Quad”

  21. By the way, the article complains that the Quad G5 is too expensive (probably why it scored 4.5/5), when there is no head-to-head price comparisons in the article. Maybe the print version has such a comparison, but the linked article doesn’t, except for a table at the very end that compares the Quad with the iMac (???) and an Intel-Dell machine that has 25% of the memory, 60% of hard-drive capacity, a weaker GPU, and a smaller monitor…

    How in the world did they conclude it was too expensive?

  22. Quoting from Jay… “Buying memory is a tricky thing, and the trickery can change everytime crucial updates it’s page.”

    Jay…. I used to buy from Crucial for all of my memory needs also. Recently I’ve bought from MEMORYTOGO dot com……. they have the same lifetime guarantee and are usually about 40% less that Crucial, which is saying a lot. Bought for 4 computers over 2 years now, no problems whatsoever.

  23. The following would be as close to a fair price comparison as I could come up with based on Canadian websites of Dell and Apple. (the reason I quote Dell’s educational prices is that you need the health care or educational sites to configure a Precision workstation).

    Dell Precision 670 XP64 Edition: CAN$10,640 educational (before tax)
    >2 x Dual-Core Intel® Xeon Processor 2.80GHz (Quad core), 2x2MB L2 cache in each processor, Hyper-Threading feature preset to ON,
    Windows XP Professional, x64 Edition with Media
    Memory: 4GB, DDR2 SDRAM Memory, 400MHz, ECC (4 DIMMS)

  24. (For some reason my previous post was truncated.. Here it is again:

    The following would be as close to a fair price comparison as I could come up with based on Canadian websites of Dell and Apple. (the reason I quote Dell’s educational prices is that you need the health care or educational sites to configure a Precision workstation).

    Dell Precision 670 XP64 Edition: CAN$10,640educational (before tax)
    2 x Dual-Core Intel Xeon Processor 2.80GHz (Quad core), 2x2MB L2 cache in each processor, Hyper-Threading feature preset to ON,
    Windows XP Professional, x64 Edition with Media
    Memory: 4GB, DDR2 SDRAM Memory, 400MHz, ECC (4 DIMMS)

  25. Hard Drive: 2 x 500GB SATA, 7200RPM Hard Drive with 16MB DataBurst Cache without RAID,

    No Floppy Drive Included,

    Single Drive: 16XDVD+/-RW ,

    No Monitor,

    Video Card: 256MB PCIe x16 nVidia Quadro FX 3450, Dual DVI or Dual VGA or DVI + VGA,

    Speakers: Dell A525 30 Watt 2.1 three piece Stereo Speakers with Subwoofer,

    Keyboard: Entry Level, USB, No Hot Keys,

    Mouse: Dell USB 2-Button Optical Mouse with Scroll,

    No Modem,

    Dell Wireless 1450 (802.11 a/b/g) WLAN USB 2.0 Adapter,

    Support: PUB/PAD:3 YRS 4X7X24 O-S, UNY,

    Software: IE, Sonic Software, CyberLink, PowerDVD, No security subscription requested, No Productivity Suite.   

    Apple PowerMac G5 Quad: CAN$9,128 regular $8,298 educational (before tax)

    2.5GHz Quad-core PowerPC G5,

    Mac OS-X v10.4 (Tiger), includes Preview 3,

    Memory: 4GB 533 DDR2 ECC SDRAM- 4x1GB,

    Hard Drive: 2x500GB Serial ATA – 7200rpm

    No Floppy drive Single drive: 6x SuperDrive DL (DVD+R DL/DVD±RW/CD-RW) Video Card: NVIDIA GeForce 7800 GT 256 MB SDRAM No Display, DVI to VGA adapter Speakers: JBL Creature II Speakers – Aluminum Apple Keyboard, Apple Mighty Mouse No modem AirPort Extreme + Bluetooth built-in AppleCare Protection Plan for Power Mac (3yr onsite + tech support) Software: Safari 2, Mail 2, Address Book 4, iChat AV 3, iCal 2, Font Book 2, DVD Player 4.5, Xcode 2, iLife ’05, GraphicConverter, Zinio Reader, OmniGraffle and OmniOutliner

  26. Note that in the above (3 posts) comparison Powermacs have a better GPU, have better software, have DL capable DVD drive, have no need for security software, can optionally have wireless, and a slew of other advantages…

  27. Oh-my-God … I just have to read that article again:

    “The Quad G5 got the highest score we’ve ever seen on the CPU-stressing CineBench rendering test: 1,104.
    The Pentium EE840 overclocked to 3.6 GHz recently got a 667, and an Athlon 64 4800+ overclocked to 2.7 GHz scored 775. CineBench is a multithreaded app, so the more cores or threads your system can handle, the more efficiently your workload gets done.”

    (shaking head in disbelief)

    Note for clarity: Both the EE840 and the 4800+, besides being 200-1100Mhz faster, are dual core CPUs as well – the best that either Intel or AMD currently offer in a desktop that isn’t running server level stuff.

    Look, I know I’m gonna get flamed, but after results like these just keep coming up I can’t help it … Is Intel and Viiv and DRM’d Boobtube 2.0 really, I mean REALLY, worth it?

    I just don’t think so.

    Why are you shaking your head in disbelief? You seem amazed a 4 core G5 system outperformed a 2 core system?

    Were you that amazed when a quad Opteron 280 scored 1335, much higher than a quad G5?
    http://www.creativemac.com/articles/viewarticle.jsp?id=36312

    Does your jaw drop just as hard after these preliminary benchmarks showing Intel’s upcoming dual Xeon processors outperform AMD’s top of the line 280?
    http://www.anandtech.com/IT/showdoc.aspx?i=2644

    Look, the only reason you’re gonna get flamed is because you can’t read. 4 cores is greater than 2 cores. No surprise, except for you it seems. Then you base this misconception on the transition to Intel by questioning if its worth it.

    If the 1000 cinebench score really impresses you, while an Opteron 280 (that scores much higher) doesn’t invoke the same response (not to mention a link showing an Intel Xeon even faster than that 280), then I shake my own head in disbelief of your continuing insistance questioning Apple’s transition based on your own flawed logic.

  28. This is the review from MacWorld on the Quad G5 2.5 Ghz. After reading it, I decided to save some money and get the G5 2.7 Ghz (the previous model), as the benchmarks were just not that spectacular on the Quad. So if anyone can decipher the disparity between the two rags, I’m all ears. I plan to make a purchase, but after MW’s review, and now this, I have no idea what to get. 🙁

    http://www.macworld.com/2005/11/reviews/quadreview/index.php

    Quote from review: “If you give just a quick glance at the Macworld Lab benchmark results for the Quad, you might wonder what the fuss is all about—on the Speedmark test suite, the new system barely managed to edge out the previous Mac performance champ, the 2.7GHz dual-processor model, and considering that the Quad has twice as much raw processing power as its single-chip, dual-core siblings, its lead over them is surprisingly modest.”

  29. Jamie:

    My quick and dirty take – The tests that shine for the Quad (Cinema 4D XL 9.1, Compressor 2.0) are the only ones that I know are definitely optimized to take advantage of more than two processors, and are processor centric in nature (as opposed to Unreal tournament, which depends a lot on how good the video card is).

    On the otherhand, while we know that iTunes, iMovie, and Photoshop are ‘multiprocessor aware’, we don’t know just how much that awareness translates into advantage. I always thought Photoshop could utilize as many CPUs as were available, but if so then it should definitely have come up with better results than this. So, I can only conclude that the tests done at MacWorld just won’t show a machine like the Quad to it’s best advantage. Maybe they’re trying to help push acceptance of the Macintel transition or something ” width=”19″ height=”19″ alt=”wink” style=”border:0;” />

    Nonetheless, the future IS with software that will take advantage of all processors available, as that’s where the entire PC industry is moving. So buying a Quad certainly won’t be a bad investment – in fact, it should only get better from here, with more and more software being optimized for it’s capabilities.

    MDN magic word – “hot”: As in the Quad is …

  30. OK sammy – I have some time today so here’s my response:

    sammy says: “Were you that amazed when a quad Opteron 280 scored 1335, much higher than a quad G5?”

    No, because I think the Opteron is a great CPU. However, while the AMD Cinebench scores on this link are nice, they aren’t a blowout, and are about on par with the advantage an integrated memory controller might give. Frankly, I’d love it if Apple was going to AMD, but that’s not the case. They’re going with Intel, who’s 2 CPU Xeon rig, with hyperthreading (and over a 1Ghz speed advantage), was beaten handily on this same test.

    you say: “Does your jaw drop just as hard after these preliminary benchmarks showing Intel’s upcoming dual Xeon processors outperform AMD’s top of the line 280?”

    No, because these results are middling at best, given the advantages that the Dempsey core – a next gen CPU – has over the present gen Opteron it’s pitted against. Bensley is a 65nm CPU (Opteron is 90nm), has twice the cache memory (4MB vs 2MB), and the chipset it’s sitting on has a faster system bus (1066 MHz vs ??? – anand doesn’t say, but I’ve never heard of an AMD FSB clocked that high) and main memory (533Mhz vs 400Mhz).

    And what does all that get the Intel CPU? A 14% faster score on an SQL test that it usually wins by 10% anyway. From Anand: “The 1066MHz front-side bus, no doubt, was required to achieve this result, along with the [higher] cache…” The Order Entry Stress Test was a statistical tie (w/in 3%), with Anand saying “The new … chipset architecture, increased front-side bus and memory bandwidth all played a part in Bensley showing this kind of improvement.” And ALL the power consumption tests resulted in Opteron victories – how does THAT look good for a 65nm CPU, or a company touting ‘performasnce per watt’ as it’s main strength? Again, from Anand, “[With] 40 servers at a datacenter, with the same power characteristics. Over a year, it would cost you $8,160 more to run the Bensley system.”

    Yeah – real impressive sam (sarcasm). Even Anandtec, which tends to be ‘pro-Intel’ in their final summations of these tests, can’t manage much more than tepid praise for the ‘new Xeon’ here.

    sammy: “… the only reason you’re gonna get flamed is because you can’t read. 4 cores is greater than 2 cores. No surprise, except for you it seems… [blah, blah, fud] … I shake my own head in disbelief of your continuing insistance questioning Apple’s transition based on your own flawed logic.”

    From the article:
    “The Quad G5 got the highest score we’ve ever seen on the CPU-stressing CineBench rendering test: 1,104. The Pentium EE840 overclocked to 3.6 GHz recently got a 667, and an Athlon 64 4800+ overclocked to 2.7 GHz scored 775.”

    My logic is as follows; these Cinebench scores show a 4 core system anywhere from about 1/3 to 1/4 faster than dualcore reference systems. Since everybody knows that multiplying cores does not result in symetrical performance gains (that is, 2x the cores never equals 2x the performance), I see the G5’s showing on that basis alone to be more than impressive. And there are other pretty obvious factors are involved that make the G5 look even better. Pentium has an over 1Ghz advantage AND hyperthreading (which should give it an effective 4 core parity vs G5 on a test like Cinebench). The latter essentially makes the competition here an ‘apples-to-apples’ comparison, and a damning one for Intel. The Opteron is also overclocked, with a 200Mhz advantage, and has an integrated memory controller which the G5 lacks. So while I give all due consideration for the great parts AMD makes, the fact is that it’s margin of loss here would be on par with a 2 core deficit. Obviously, a better test would have been a single dualcore G5 vs a single dualcore Opteron280 vs Pentium EE840, but until then these results will have to serve as the signposts. And it doesn’t take eyeglasses to see that the G5 is ahead of the curve.

    I get flamed by people like you because your stuck so far up Jobs’ butt you’ve forgotten what sunshine smells like. I’ve been right about how video drove this transition, and it drives you fanbois crazy. The preformance numbers keep bareing out that Intel is not doing all that well, despite their hype, and if Apple was concerned about performance as they say, they either would have stood pat with PPC or went with AMD. You’re a smart guy Sammy, so the fact that you allow yourself to also NOT see what’s right in front of your face is another ‘head shaker’ for me.

    But just one among many, I’m afraid. ” width=”19″ height=”19″ alt=”oh oh” style=”border:0;” />

  31. Yeah – real impressive sam (sarcasm). Even Anandtec, which tends to be ‘pro-Intel’ in their final summations of these tests, can’t manage much more than tepid praise for the ‘new Xeon’ here.

    And taking into consideration the points you outline above, you still refused to give Intel their due concerning their upcoming Yonah processor which outperforms AMD clock for clock while utilizing considerably less power. You ignored this link prior but I will bring it up yet again, as it addresses your “concerns” directly in your attempt above to dismiss the Bensley architecture:

    … We continue to see that the Core Duo can offer, clock for clock, overall performance identical to that of AMD’s Athlon 64 X2 – without the use of an on-die memory controller. The only remaining exception at this point appears to be 3D games, where the Athlon 64 X2 continues to do quite well, most likely due to its on-die memory controller….
    http://www.anandtech.com/cpuchipsets/showdoc.aspx?i=2648&p=14

    Pentium has an over 1Ghz advantage AND hyperthreading (which should give it an effective 4 core parity vs G5 on a test like Cinebench). The latter essentially makes the competition here an ‘apples-to-apples’ comparison, and a damning one for Intel.

    First of all the 1Ghz clock advantage means absolutely nothing (and it should be obvious to you as well) simply because of the deeply pipelined nature of Netburst. For you to even suggest it just tells me you simply don’t understand the underlying technology (yet again). A deeper staged pipeline allows less processing per cycle (compared to AMD and PowerPC) but much higher clockspeeds to compensate. Therefore, a clockspeed of a 20 staged pipeline CPU cannot be directly compared to a 14 staged pipeline CPU.

    The 20-stage pipeline on the Pentium 4 is what allows it to hit higher clock speeds right off the bat without requiring a die shrink. It is for this reason that the Pentium 4 will debut at speeds of 1.4GHz and higher (we will talk more about clock speed in a bit). Before you let that number impress you too much, you have to realize that the 20-stage pipeline of the Pentium 4 also yields what is called a lower amount of Instructions Per Clock (IPC). A lower IPC basically means that you get less accomplished in a given amount of time when compared to a processor that has a higher IPC – pretty simple right?
    http://www.anandtech.com/showdoc.aspx?i=1301&p=3

    Secondly, you simply blow me away with your statement that hyperthreading allows the Pentium EE to compete “apples to apples” with the Quad G5. How ignorant are you Odyssey? Do you seriously believe a dual core Pentium EE is really 4 cores in disguise? Do you know what hyperthreading is?

    Hyperthreading is the ability to process secondary threads on unused CPU cycles. Obviously, if a core is utilized 100%, how is there room for parallel threads to process?.

    Programmers have long known that some applications will run more efficiently if they’re coded into a series parallel tasks, called threads. Modern multi-processing operating systems can then schedule those threads to operate on each of a system’s two or more CPUs, just as it schedules the applications and other processes themselves.

    Intel’s technology essentially fools the operating system into thinking it’s hooked up to two processors, allowing two threads to be run in parallel, both on separate ‘logical’ processors within the same physical processor. The OS sees double through a mix of shared, replicated and partitioned chip resources, such as registers, maths units and cache memory.

    According to Intel, less that five per cent of the Xeon’s die area is taken up the SMT-enabling circuitry, primarily because most of the functionality is provided by chip components that would otherwise be standing idle. The chip maker estimates that a single instruction thread only uses around 35 per cent of a processor’s available resources. Running a second thread allows those otherwise idle circuits to be do some work.

    So if one thread is busily hacking away at a list of integer values, the floating-point units are free to crunch numbers for a second thread. Such a smooth division of labour is uncommon, alas, so the best HT can do is increase certain applications’ performance by up to 30 per cent, according to Intel, though it admits the average gain is more like 10-20 per cent.

    That clearly isn’t going to allow a single Xeon MP to match the performance of a HT-less multi-processor rig, but it does provide a significant boost.
    http://www.theregister.co.uk/2002/06/18/what_the_hell_is_hyperthreading/

    Now tell me again, how is it that you see a dual Pentium as a “true 4 core processor being directly comparable to a Quad G5”?

  32. Obviously, a better test would have been a single dualcore G5 vs a single dualcore Opteron280 vs Pentium EE840, but until then these results will have to serve as the signposts. And it doesn’t take eyeglasses to see that the G5 is ahead of the curve.

    Where’d you get those glasses?

    http://www.digitalproducer.com/articles/viewarticle.jsp?id=32951-1&afterinter=true

    Granted, that is not a true dual core G5 system, however perform would be quite nearly the same…and not “ahead of the curve” as you put it.

  33. Sammy says: “you still refused to give Intel their due concerning their upcoming Yonah processor which outperforms AMD clock for clock while utilizing considerably less power.”

    Yonah wasn’t being tested here – I don’t have the same penchant for going off topic as you – but if you want to bring it up, fine. Its GREAT that Intel has finally made a CPU that can compete with AMD offerings. The issue that you keep ignoring is that this is Intels ‘next gen’ technology, and AMD still runs neck and neck with it. When AMD introduces it’s 65nm CPUs in 06, and Yonah is still there, then talk to me. Until then, this is just a momentary blip for Intel, based on timing, not technical prowess.

    sammy:”1Ghz clock advantage means absolutely nothing (and it should be obvious to you as well) simply because of the deeply pipelined nature of Netburst. For you to even suggest it just tells me you simply don’t understand the underlying technology (yet again). A deeper staged pipeline allows less processing per cycle (compared to AMD and PowerPC) but much higher clockspeeds to compensate. Therefore, a clockspeed of a 20 staged pipeline CPU cannot be directly compared to a 14 staged pipeline CPU.”

    Spoken like a true Pot. Call the Kettle “black” much? Only an idiot would claim that 1Ghz means “nothing”, and I’m not going to bother being drug into that bit of nonsense. And in your little tutorial on pipelines, you neglect to even recognize that those extra 6 stages Intel has are SUPPOSED to yield more work being done, not just higher clockspeeds. The fact that they don’t is a function of Intel’s lousy branch prediction technology, yielding miss-hits that have to flush all the info out of the pipeline in a given clock tic, and start the process ALL OVER AGAIN. This is why Intel’s 1Ghz advanatage yields them nothing; b/c they get less work done in 3+Ghz than everyone else does in 2+Ghz. Yonah is a retreat back to more a manageable pipeline size for them, which also happens to save them the trouble of dealing with all the speed induced heat issues Pentiums have had up to now, but the fact remains that it took them 3 years to figure this out. AMD meanwhile continues to refine an already solid design. That you seem to be ignorant of that tells me that the only fool here is you.

    sam:”… you simply blow me away with your statement …”

    Yeah, I’ve noticed that facts seem to have that effect on you.

    “… that hyperthreading allows the Pentium EE to compete “apples to apples” with the Quad G5. How ignorant are you Odyssey? Do you seriously believe a dual core Pentium EE is really 4 cores in disguise? Do you know what hyperthreading is? Hyperthreading is the ability to process secondary threads on unused CPU cycles. Obviously, if a core is utilized 100%, how is there room for parallel threads to process?.”

    You’re funny. Name calling is clearly what you’re best at, but you’re penchant for ‘echo chamber’ tutorials, when it’s you who need to brush up, is a close second. It would be comical, if you didn’t pass yourself off as an ‘expert’.

    Lessee – from Intel we get the following: “The Intel® Pentium® processor Extreme Edition combines HT Technology with dual-core processing to give people PCs capable of handling four software threads. HT Technology enables gaming enthusiasts to play the latest titles and experience ultrarealistic effects and gameplay. And multimedia enthusiasts can create, edit, and encode graphically intensive files while running a virus scan in the background.”
    http://www.intel.com/technology/hyperthread/

    Hey, if you can’t handle the truth from me, take it from the source. Notice the words “… combines HT Technology with dual-core processing to give people PCs capable of handling four software threads”? You do realize that a four core system without hyperthreading would also handle four software threads? Get the connection now? sheesh.

    cont …

  34. Now, I did say that “essentially” it was an apples-to-apples comparison (time for a grammer tutorial), which is differnt from saying “completely”, and the reason I did that is – yes Virginia – there IS a difference between HT and 4 real cores. But the fact is that with software that utilizes it, like CineBench, HT is a demonstratable performance booster for Intel (one of the few), so for the EE to have done this poorly against the Quad here can not be overlooked. Nice try, but no kewpie doll for you.

    s-guy:”Where’d you get those glasses?”

    They’re X-Ray Specs, from the back of my Archie comic book. Yours?

    “http://www.digitalproducer.com/articles/viewarticle.jsp?id=32951-1&afterinter=true
    Granted, that is not a true dual core G5 system, however perform would be quite nearly the same…and not “ahead of the curve” as you put it.”

    Amazing how forgiving you are with your own examples. The fact that it’s not a “true dual core system” is the horsefly in your ointment. Plus, the memory latencies that the last gen PowerMac chipset was saddled with, and the performance penalty this inevitably imposes against a true dual core CPU with an INTEGRATED MEMORY CONTROLLER like AMD’s, makes this example of yours as an indicator of CPU capability nearly worthless.

    Here we see the Quad on many of the same tests:
    http://www.barefeats.com/macvpc.html

    Opteron and Quad are neck and neck (even though the G5 still doesn’t have an integrated memory controller), and the Mac stomps the Intel offerings on all but the Maya render tests, which themselves bring the GPU in to play. “… a new trend [is] to involve the graphics card in what has been traditionally a CPU only function. Two examples of applications that involve the GPU are Maya and Motion.”

    Two tests, two very different results, but … hmmm. Only the latter deals with the Mac that we were originally supposed to be talking about. Niether is directly transferable to the other, but you obviously brought in a test suite ‘ringer’ in order to bolster your argument. Ah well. In any event, taken together they both demonstrate what I’ve been saying since day one – that AMD CPUs are the best x86s available, and thus anyone claiming the move to Intel was made for performance reasons is being willfully stupid.

    Apparently, this includes you sammy.

  35. Yonah wasn’t being tested here – I don’t have the same penchant for going off topic as you – but if you want to bring it up, fine. Its GREAT that Intel has finally made a CPU that can compete with AMD offerings. The issue that you keep ignoring is that this is Intels ‘next gen’ technology, and AMD still runs neck and neck with it. When AMD introduces it’s 65nm CPUs in 06, and Yonah is still there, then talk to me. Until then, this is just a momentary blip for Intel, based on timing, not technical prowess.

    Again spoken like a true biased, clueless person who “pretends” to be knowledgeable in a subject beyond his own capability of comprehension. Intel’s 65nm chips are weeks away from debut…..AMD’s will not have a 65nm offering until well into 2006. In fact, the only upcoming releases from AMD in the near future is a higher clocked dual core FX-60 chip and a speed bumped Opteron. Your penchant for lambasting Intel on upcoming technologies in comparison to current rival tech, all the while praising Cell in striking contrast of the very same arguments you presented, shows just what some biased moron you are on some agenda….which everyone clearly sees.

    Let’s see that statement once again: When AMD introduces it’s 65nm CPUs in 06, and Yonah is still there, then talk to me. Until then, this is just a momentary blip for Intel, based on timing, not technical prowess..

    Ok moron, so you’re thrashing Intel for bringing their next gen technology earlier than the competition? (Which is already in the hands of reviewers and acknowledged for its clock for clock superiority over AMD’s current best)Then your dumbass go on to praise “future” Cell variants that’s not even in existence yet, get uber-excited over in house IBM/Sony selective comparisons and press releases, all while completly ignoring the outlash from developers with actual development kits.

    And then to show just how clueless you are to everyone here, you publicly portray your “amazement” of the quad G5 cinebench scores over dual core CPU’s.

    You can’t possibly be that dense, can you Odyssey?

    (more to come).

  36. odyssey67, this is directed to you.

    when sammy FINALLY posted that, of course, a 4 core pmac is going to out perform a 2 core system from both AMD and Intel, why did you post a bunch of crap about nothing pretty much. you did direct your post to sammy, but why post a bunch a of garbage when the original arguement was simply about how the benchmark was unfairly setup.

    of course a 4 quad pmac gets a ~1100 on a cinebench, while an AMD 2.7 dualcore gets 778. But go quad AMD’s, they now give you ~1300. hrm almost 2x the performance actually….

  37. Spoken like a true Pot. Call the Kettle “black” much? Only an idiot would claim that 1Ghz means “nothing”, and I’m not going to bother being drug into that bit of nonsense. And in your little tutorial on pipelines, you neglect to even recognize that those extra 6 stages Intel has are SUPPOSED to yield more work being done, not just higher clockspeeds. The fact that they don’t is a function of Intel’s lousy branch prediction technology, yielding miss-hits that have to flush all the info out of the pipeline in a given clock tic, and start the process ALL OVER AGAIN. This is why Intel’s 1Ghz advanatage yields them nothing; b/c they get less work done in 3+Ghz than everyone else does in 2+Ghz.

    YET AGAIN….YET AGAIN….YET AGAIN….your utter stupidity in the technology being discussed is highly apparent. The extra 6 stages was NEVER “supposed” to yeild more work being done for each clock cycle….did you even bother to read the link I provided for you? It has NOTHING to do with a “lousy” branch predictions…in fact, please explain what EXACTLY is lousy about Intel’s branch prediction algorithm..

    A longer staged pipeline will ALWAYS be less efficient than a shorter stage pipeline in terms of branch prediction, that is simple EE logic (10% missed prediction on an 11 stage CPU is less significant than a 20 stage CPU). Clock speed becomes less indirectly comparable just because of that fact (just as the article states, a 1Ghz Pentium 3 is faster than a 1Ghz Pentium 4). However, having a smaller, longer stage pipeline allowed for much higher clockspeed to compensate. In contrast, an Opteron is a 12 stage pipelined CPU…the fact that it cannot (currently) hit beyond 3Ghz is not a coincidence.

    (more to come)

  38. Hey, if you can’t handle the truth from me, take it from the source. Notice the words “… combines HT Technology with dual-core processing to give people PCs capable of handling four software threads”? You do realize that a four core system without hyperthreading would also handle four software threads? Get the connection now? sheesh.

    No, the question is do YOU understand. A dual core HT processor can of course handle 4 simultaneoous threads, in certain conditions ONLY and NEVER with full processor speed. No surprise the customer relations quote you copied from their website naturally excluded those limitations, as your brain seems to be able to only comprehend general, layman and customer orientated quick quotes rather than deeply detailed, technical overviews.

    Hyper-threading’s greatest strength–shared resources–also turns out to be its greatest weakness, as well. Problems arise when one thread monopolizes a crucial resource, like the floating-point unit, and in doing so starves the other thread and causes it to stall. The problem here is the exact same problem that we discussed with cooperative multi-tasking: one resource hog can ruin things for everyone else. Like a cooperative multitasking OS, the Xeon for the most part depends on each thread to play nicely and to refrain from monopolizing any of its shared resources.

    For example, if two floating-point intensive threads are trying to execute a long series of complex, multi-cycle floating-point instructions on the same physical processor, then depending on the activity of the scheduler and the composition of the scheduling queue one of the threads could potentially tie up the floating-point unit while the other thread stalls until one of its instructions can make it out of the scheduling queue.

    http://arstechnica.com/articles/paedia/cpu/hyperthreading.ars/5

    I ask you again Odyssey, do you know what hyperthreading is? Do you see why a hyperthreaded enabled processor is NOT equavalent to two seperate processors and performance will NEVER be equal between the two? Do you understand why a dual core HT processor CANNOT process 4 simultaneous floating point threads as fast as two dual core physical cores?

    Hyperthreading allows threads to be processed utilizing unused portions of CPU cycles that are not being utilized in a primary thread (a thread that primarily uses the integer unit of a CPU frees up threads to use the floating point unit).

    You cannot make a V12 engine out of a V6. Perhaps you will understand this zipper analogy:

    The words “parallel” and “simultaneous” are somewhat misleading here since true parallelism would include that everything is processed simultaneously through two parallel pipes. A more appropriate picture would be a zipper where data chunks from two parallel instruction sets are processed in alternating chunks by interleaving data and instructions. On a processor level, this would mean that the CPU continuously switches from one application to the other in a few nanoseconds intervals. The problems associated with this technology are that additional buffers are required for holding the data internally and for keeping record of the status of each set of instructions.
    http://www.lostcircuits.com/cpu/p4_306/

    This is by the far the biggest ignorance you have shown us so far, eclipsing the little misunderstandings you had about Maya, OpenGL, Cell, etc.. (jeez, the list is growing, isn’t it?)

    Once again, a 4 core system handling 4 threads is NOT equavalent to a 4 threads handled under a single dual core hyper threaded processor.

    (more to come)

  39. Regardless piplines, branch predictions, mask sizes, power consumption, integrated memory controllers, etc., etc. the fact is:

    – 4 cores barely besting 2 cores isn’t much cause for celebration.

    Personally, regardless of why the PPC may be great and the Opteron nice, etc. it’s the system that’s being tested. Bottom line, if I need to do a bunch of rendering, photo processing, etc., which allows me to get more done?…. And with similarly configured “accessories” (like GPUs, memory from same manufacturer, hard drives, etc.) what’s the cost?

  40. Now, I did say that “essentially” it was an apples-to-apples comparison (time for a grammer tutorial), which is differnt from saying “completely”, and the reason I did that is – yes Virginia – there IS a difference between HT and 4 real cores. But the fact is that with software that utilizes it, like CineBench, HT is a demonstratable performance booster for Intel (one of the few), so for the EE to have done this poorly against the Quad here can not be overlooked. Nice try, but no kewpie doll for you.

    There is a performance boost in most cases with HT enabled, however not as dramatic as you are claiming…..your insistance that a 4 physical core system is comparable to 4 logical cores on a two physical core system is completely off base. And here’s the proof:

    Cinebench 2003:
    Pentium 4 3.2 w/o HT: 323
    Pentium 4 3.2 WITH HT: 381

    Xeon 2.8 Ghz single CPU: 279
    Xeon 2.8 Ghz dual CPU: 519

    BIG difference there Odyssey, between two logical cores and two physical cores (and with a slower clockspeed no less!).

    Now, do you still think a Pentium EE’s logical 4 cores should produce an equavalent cinebench score of the Mac’s physical 4 scores?

    Amazing how forgiving you are with your own examples. The fact that it’s not a “true dual core system” is the horsefly in your ointment. Plus, the memory latencies that the last gen PowerMac chipset was saddled with, and the performance penalty this inevitably imposes against a true dual core CPU with an INTEGRATED MEMORY CONTROLLER like AMD’s, makes this example of yours as an indicator of CPU capability nearly worthless.

    You are actually correct in asserting this is not a valid comparison in dual core vs dual core. However, there is nothing to remotely suggest the dual core G5 would be “ahead” of the curve in any dual core vs dual core comparison, as the purpose of the links shows…. the Pentium’s 715 cinebench score vs 581 for the dual 2.5 G5, not the AMD score.

    Here we see the Quad on many of the same tests:
    http://www.barefeats.com/macvpc.html

    Opteron and Quad are neck and neck (even though the G5 still doesn’t have an integrated memory controller), and the Mac stomps the Intel offerings on all but the Maya render tests, which themselves bring the GPU in to play. “… a new trend [is] to involve the graphics card in what has been traditionally a CPU only function. Two examples of applications that involve the GPU are Maya and Motion.”

    Excuse me, where did you see a Quad in that test? All I see are dual configuration single core CPU’s. The Opterons tested there aren’t even the newer revision E chips with SSE3/DDR2 support. Are you confusing yourself Odyssey?

    And then you dismiss the Maya benchmark as it is due to the graphics card being “in play”….what’s the matter, those glasses of yours don’t allow you to see the software results, where the video card do not come into play at all?

    END

  41. Holy Crap! … What was I thinking, coming back to this thread?? I don’t know how to convey, via keyboard, the big belly laugh I got from seeing all this! Well, I doubt I’ll equal the quantity of your output (maybe), but I’ll try to address the pertinient points.

    Sammy says: “Cinebench 2003: …
    … Now, do you still think a Pentium EE’s logical 4 cores should produce an equivalent Cinebench score of the Mac’s physical 4 scores?”

    Typical that I have to point this out AGAIN, but I never claimed absolute equivalency, and I made the distinction for many of the reasons you mention. My ‘oversight’, apparently, was in giving you credit for understanding the difference’s involved – and me the break of writing a term paper about ’em. You spending ungodly amounts of time & space setting up the straw man, that I said what I didn’t, is ‘Pure Sammy’ – your need to impress yourself clearly knows few bounds, but there’s no mistaking it.

    Anyway, in answer to your question – here’s the author of the Ars article you linked to: “As bus speeds increase, and more cache becomes available on die, hyper-threading is … more and more efficient. It appears to be somewhat of an engineering symbiotic relationship.”

    These test numbers you’ve provided are from older Intel CPUs, that are down in both catagories when compared to what PCMag tested (2MB cache and 800MHz FSB).
    http://www1.us.dell.com/content/topics/topic.aspx/global/products/dimen/topics/en/dimen_xps600_sp_specs?c=us&cs=19&l=en&s=dhs

    So, while I’ve read Ars here, frequently, unlike you I don’t only incorporate SOME parts of the article into my thinking. Your zipper analogy is apt for understanding why I see present day benchmarks vs. versions from 2-3 yrs ago as invalid. Again, the Ars article: “Like a cooperative multitasking OS, the [HT] Xeon for the most part depends on each thread to play nicely and to refrain from monopolizing … its shared resources… In sum, resource contention is definitely … why … With the wrong mix of code, hyper-threading decreases performance, just like it can INCREASE performance with the RIGHT mix of code.” Translation: If the software isn’t optimized to deal with HT the zipper gets bunged up, but if it IS optimized it just zips.

    Now, I’ll leave it to you to dig up the links, but I’ll bet dollars to donuts that Cinebench 05 (in PCMag) does a better job of ‘zipping’ w/HT than the older versions from two years ago (2004-01-07) did. It’s certainly common for these tests to optimize for new CPU resources as they become more prevalent. So, while HT was pretty new back then, to the extent that it’s limitations & strengths can be accomodated & leveraged in the test, I’m sure they have been.

    Repeat myself time: I don’t think this means parity with distinct cores, especially across the board. But with a performance specific app like Cinebench, and with the modern cache-heavy CPUs Intel has now, HT will be a real boost – much more than these ancient results indicate. So, while I didn’t invent the thing, the only one showing they don’t understand Hyperthreading here is you. And it’s a self imposed ignorance – the worst kind.

    Regarding the Barefeets test, you ask: “Excuse me, where did you see a Quad in that test? … [and] The Opterons tested … aren’t even the newer revision E chips with SSE3/DDR2 support. Are you confusing yourself Odyssey?”

    Yeah, I was. But, fortunately, it’s a confusion that benefits my argument. Even though the G5s in this test are single core, so are the Opterons, & they’re both – yet again – one-two on the chart, with Intel’s very best huffing along (a full Gig faster) just to keep pace. These tests simply prove my point with last gen technology: For the money, Intel is average, PPC is better, AMD may be even better still; thus Apple going to Intel was clearly a move to get their DRM and chipset gimmicks for the video device/content market, and not for performance.

    Oh, and you’re swinging wild again here too. Which Opteron are we talking about? The 252? Here:
    “The Opteron 252 processor runs at 2.6 GHz, the same [as] … the Athlon64 FX-55 [although]… The 252’s 90nm manufacturing and support for SSE-3 make this chip more advanced … The Opteron 252 has 128k of L1 and 1 MB … of L2 cache, which is un-changed from its original design… [but] still has an on-die dual channel DDR-400 memory controller …”
    http://www.gamepc.com/labs/view_content.asp?id=x36o252&page=3&MSCSProfile=95385A1F52DEA1A229D5B37542054464256B554C9D0CCF491BFEC447475DF167A3AE26114C22198F86C9DAC127C814C83FDFC7284411D68332F7F1D343AE9DE1AC2D2EB3DC6DA2ED049B02645B2E68EDDEF4B900C9B308F48D567E8907197391A0834FF6FA617E6C68C8F379867AD272A0B632EB323B98FE08783E001ABDE9930A5AEBC28D2EB11D

    So, aside from no DDR2 support – which is inconsequential anyway since the G5 didn’t have it then either – you got almost everything wrong. Good job!

    cont…

  42. Sammy says: “YET AGAIN….YET AGAIN….YET AGAIN….your utter stupidity in the technology being discussed is highly apparent.”

    Calm down you goofball. The only thing that’s “highly apparent” is your short fuse. The universal mark of immaturity.

    Sammy: “The extra 6 stages was NEVER “supposed” to yeild more work being done for each clock cycle….did you even bother to read the link I provided for you?”

    I read it, AND IT’S WRONG. Here:
    “In general, the speedup in completion rate versus a single-cycle implementation that’s gained from pipelining is ideally equal to the number of pipeline stages. A four-stage pipeline yields a four-fold speedup in the completion rate versus single-cycle, a five-stage pipeline yields a five-fold speedup, a twelve-stage pipeline yields a twelve-fold speedup, and so on. This speedup is possible because the more pipeline stages there are in a processor, the more instructions the processor can work on simultaneously and the more instructions it can complete in a given period of time. So the more finely you can slice those four phases of the instruction’s lifecycle, the more of the hardware that’s used to implement those phases you can put to work at any given moment.”
    http://arstechnica.com/articles/paedia/cpu/pipelining-2.ars/1

    I urge you to read the rest of the article for yourself – even you might learn something.

    Sammy: “It has NOTHING to do with a “lousy” branch predictions…”

    Here’s what your own link says: “Modern day CPUs attempt to increase the efficiency of their pipelines by predicting what they will be asked to do next. This is a simplified explanation of the term Branch Tree Prediction. When a processor predicts correctly, everything goes according to plan but when an incorrect prediction is made, the processing cycle must start all over at the beginning of the pipeline. Because of this, a processor with a 10 stage pipeline has a lower penalty for a mis-predicted branch than that of a processor with a 20 stage pipeline. The longer the pipeline, the further back in the process you have to start over in order to make up for a mis-predicted branch.”

    Who’s having the problem with reading things again?
    ” width=”19″ height=”19″ alt=”hmmm” style=”border:0;” />

    Sam: “in fact, please explain what EXACTLY is lousy about Intel’s branch prediction algorithm…”

    What’s lousy about it is that it’s not up to the job of dealing with longer pipelines! I mean, you put the same technology in a shorter pipelined CPU (as is happening with Yonah) then the miss-hits don’t impose as much of a penalty. But, as per Intel practice for the last few years, they didn’t sweat the details when they decided to go with 20 staged Netburst. They did address the problem later, with the Prescott revisions, but in addition to beefing up the branch prediction they lengthened the pipeline AGAIN, to 30 stages. That kept it on pace with the older Northwood version, but AMD still beat it on Winstone tests (which are branchy in nature). Here:
    “In both tests, Prescott is essentially in a statistical dead heat with the 3.2GHz Northwood part — slightly behind in Business Winstone and slightly ahead in CC Winstone. Given the branchy nature of business applications, Prescott’s performance in Business Winstone is surprisingly good. Clearly the larger cache, improved branch prediction and better memory handling offsets the deeper pipeline.”
    http://www.extremetech.com/article2/0,1697,1606752,00.asp

  43. (Yikes! Don’t know what happpened with the looong lines there … must have been the link. ah well.)

    Moving on- the most complete article I’ve read on branch prediction is found here:
    http://arstechnica.com/articles/paedia/cpu/pentium-m.ars/1

    It delves into the Pentium M’s improvements, Pentium4s shortcomings, and more. Specifically, it mentions 970s (G5s) clear advantages here, as well as in simply not having to deal with other legacy overhead that weighs down Intel. I think – IF YOU READ IT – you should have no more trouble understanding my statement regarding Intel’s ongoing struggles with branch prediction. I will say that the future for them looks brighter on that score (mostly because they are finally abandoning Netburst), but they have a lot of ground to make up, AND their competitors aren’t sitting still either.

    Sammy: “A longer staged pipeline will ALWAYS be less efficient than a shorter stage pipeline in terms of branch prediction, that is simple EE logic (10% missed prediction on an 11 stage CPU is less significant than a 20 stage CPU). Clock speed becomes less indirectly comparable just because of that fact (just as the article states, a 1Ghz Pentium 3 is faster than a 1Ghz Pentium 4). However, having a smaller, longer stage pipeline allowed for much higher clockspeed to compensate. In contrast, an Opteron is a 12 stage pipelined CPU…the fact that it cannot (currently) hit beyond 3Ghz is not a coincidence.”

    What the hell’s a “smaller, longer stage pipeline” supposed to mean?? It’s one or the other genius.

    A longer staged pipeline WILL NOT “always be less efficient.” It can be, if (all other things being equal) A] the entire pipeline is filled to capacity only sporadically, and B] that capacity is consistently wasted by a branch prediction miss-hit. That’s when your percentages example comes into play. The two conditions are mutually exclusive though. In other words, filling the pipeline to capacity would, by itself, have an adverse impact on performance (regardless of the accuracy of branch prediction algorithms) just b/c it intuitively takes longer to fill anything to capacity. On the other hand, a deeply pipelined processor can overcome this problem if it goes for long periods with its pipeline full, essentially by leveraging it’s increased clockspeed to overcome the slower fillrate per clock cycle. And, obviously, if a pipeline is not filled to capacity it won’t make much difference if a miss-hit occurs. However, if both happen during the same cycle, then your looking at the worst of both worlds. Yes, a longer pipeline will impose a bigger cost if all things don’t go well, but on it’s face that’s not a forgone conclusion, and the other advantages a longer pipeline provides could – if executed well – equalize or overcome the potential pitfalls. Again, go to the Ars article I linked.

    As an aside, it’s my opinion that this badly executed Netburst architecture is probably the reason why Intel developed HyperThreading in the first place.
    Again from Ars: “SMT not only improves application performance by increasing a multithreaded application’s average completion rate under normal circumstances (i.e. all instructions are found in the L1 cache), but it also can prevent the completion rate from dropping to zero as a result of cache misses and memory latencies. When the processor is executing two threads simultaneously, and one of the threads stalls in the fetch stage (i.e., there’s a cache miss so the thread must fetch instructions or data from main memory), the processor can continue normally executing the non-stalled thread.” A “cache miss” here is referring to a branch miss-predict.
    http://arstechnica.com/articles/paedia/cpu/pipelining-2.ars/5

    There apparently wasn’t too much they could do about the branch prediction problem, so they devised HT as the cheapest way of keeping the CPU working even when a miss-hit did occur. But that’s just my take.

    BTW – I never claimed to be an expert Sammy. That kind of self-aggrandizment is definitely more your thing. In point of fact, I am – like you – simply an every-day enthusiast, and proud of it. The difference is that I actually bother to read entire articles written by people who are experts, and incorporate ALL their information. You see little snippets, match them up with your specious arguments against whoever you happen to be torked at, and off you run – linking like a crazy-man, boldfacing and all-caps all over the place, and clearly getting a hard on from calling me ‘stupid’ and so forth. It’s hysterical, especially considering that you’ve demonstrated pretty spotty mental horsepower on this topic yourself.

    You’re a living testament to the old adage: Talk is indeed cheap.

  44. Typical that I have to point this out AGAIN, but I never claimed absolute equivalency

    And there-in lies the biggest “belly laugh” of all. You’re lucky you decided to reply when this thread is now buried in archives, as you’ve saved grace from further people reading first hand of your technical ignorance and stupidity. However, that didn’t stop your initial post from becoming the laughing stock of this thread…..your absolute “amazement” and “jaw dropping” response to the quad G5 scores in relation to the dual core Pentium and Opteron was quite a laugh for everyone who read your joke of a response. And NOW, you try to downplay your idiocy that you never expected “absolute equavalency” in a vain attempt to hide your ignorance. These are your own words: …AND hyperthreading (which should give it an effective 4 core parity vs G5 on a test like Cinebench). The latter essentially makes the competition here an ‘apples-to-apples’ comparison, and a damning one for Intel.. You can’t hide from your ignorance Odyssey…..EVERYONE[/B] here sees it, now matter how hard you try to spin it.

    These test numbers you’ve provided are from older Intel CPUs, that are down in both catagories when compared to what PCMag tested (2MB cache and 800MHz FSB).

    The fact that older CPU’s were used doesn’t change in any way how hyperthreading works, which hasn’t changed at all since its debut. Hyperthreading still shares the L1, L2 and L3 caches of one physical processor and all logical cores have to share the same physical front side bus…. that will never change. You are obviously running away now from the facts, since when faced with proof you have absolutely nothing to offer. An increase in cache and front side bus still isn’t going to make 4 logical cores any closer to 4 physical cores. And nice link that doesn’t work by the way.

    So, while I’ve read Ars here, frequently, unlike you I don’t only incorporate SOME parts of the article into my thinking. Your zipper analogy is apt for understanding why I see present day benchmarks vs. versions from 2-3 yrs ago as invalid. Again, the Ars article: “Like a cooperative multitasking OS, the [HT] Xeon for the most part depends on each thread to play nicely and to refrain from monopolizing … its shared resources… In sum, resource contention is definitely … why … With the wrong mix of code, hyper-threading decreases performance, just like it can INCREASE performance with the RIGHT mix of code.” Translation: If the software isn’t optimized to deal with HT the zipper gets bunged up, but if it IS optimized it just zips.

    When hyperthreading fails to provide a performance it is due to the nature of hyperthreading itself. For example, often when encoding turning off hyperthreading can actually become a benefit under certain encoders as the logical processors move large chunks of data around a single physical cache, creating huge overhead…as you can see here:
    http://www.2cpu.com/articles/43_3.html

    There is no “wrong mix of code” in those conditions….just the limitations of having two “fake” processors on one physical die.

    Hyperthreading is also poor for server applications for the very same shared resources reasons:
    With both SQL Server and Citrix Terminal Server installations, HT-enabled motherboards show markedly degraded performance under heavy load. Disabling HT restores expected levels, according to reports from within the IT industry.

    “Our customers observed very interesting behaviour on high-end HT-enabled hardware. They noticed that in some cases when high load is applied SQL Server CPU usage increases significantly but SQL Server performance degrades,” wrote Ocks.

    Ocks then detailed testing which showed this behaviour where a system thread — in this case one cleaning out blocks of disk cache memory — is running at the same time as worker threads. “With Intel HT technology, logical processors share L1 & L2 caches. As you would guess [this] behaviour can potentially trash L1 & L2 caches,” he said.
    http://news.zdnet.co.uk/hardware/chips/0,39020354,39237341,00.htm

    And you need to reread that Arstechna article again, as it supports my arguments fully. You haven’t shown me anything that contradicts what I’ve stated. Find me one ONE thing in that article that supports any statement of yours.

  45. Now, I’ll leave it to you to dig up the links, but I’ll bet dollars to donuts that Cinebench 05 (in PCMag) does a better job of ‘zipping’ w/HT than the older versions from two years ago (2004-01-07) did. It’s certainly common for these tests to optimize for new CPU resources as they become more prevalent. So, while HT was pretty new back then, to the extent that it’s limitations & strengths can be accomodated & leveraged in the test, I’m sure they have been.

    Leave me to dig up links? Bet me dollars to donuts? You sure are a character Odyssey, as your habit of making claims without even a shred of evidence or logical reasoning is rapidly becoming your trademark these days. Let me ask you something, do you know what Cinebench is? I mean, do you know what sort of number crunching when an scene is being rendered? Take a guess braniac….and tell me how exactly a newer version of a render would radically change a floating point/integer process. If Cinebench is heavily floating point based, the number of logical processors will NOT provide a substantial performance boost as a the processing power of the physical FPU is static. 2 logical processors is ONE physical FPU unit, NOT 2…and as I have proved to you above…two physical processors is helluva faster than two logical processors (as the number of physical FPU’s are greater, not shared under logical processor threading……no amount of hyperthreading optimizations will change that. This is why under the best conditions (i.e, a good mix of integer/floating point operations), you will see at best, 15-20% improvement.

    So YOU prove yourself here Odyssey, as I have proven myself. SHOW me the benchmarks that prove otherwise.

    Yeah, I was. But, fortunately, it’s a confusion that benefits my argument. Even though the G5s in this test are single core, so are the Opterons, & they’re both – yet again – one-two on the chart, with Intel’s very best huffing along (a full Gig faster) just to keep pace. These tests simply prove my point with last gen technology: For the money, Intel is average, PPC is better, AMD may be even better still; thus Apple going to Intel was clearly a move to get their DRM and chipset gimmicks for the video device/content market, and not for performance.

    I already explained to you the megahertz differential of Netburst cannot be directly compared. If you insist on being so stubbornly ignorant, then I can’t really help you (and, Opteron vs G5 regardless of generation always favor AMD by a good margin in countless other reviews except for the select few renderers Rob Morgon insists on running).

    And here is where your stupidity rears its ugly head, and why so many people have issues with you. You base your entire arguments on old netburst architecture….despite the fact Apple will NOT be using these chips in their transistion. The Yonah and future variants have strong public support because the technology is there to put Intel right up there in the performance throne (as the countless links I provided you from Tom’s Hardware, Anandtech, etc… clearly show). Everyone likes what Intel has on their roadmap…..EXCEPT YOU, and you’re too stupid to even realize the progress they have made from the Netburst days.

    Oh, and you’re swinging wild again here too. Which Opteron are we talking about? The 252? Here:
    “The Opteron 252 processor runs at 2.6 GHz, the same [as] … the Athlon64 FX-55 [although]… The 252’s 90nm manufacturing and support for SSE-3 make this chip more advanced … The Opteron 252 has 128k of L1 and 1 MB … of L2 cache, which is un-changed from its original design… [but] still has an on-die dual channel DDR-400 memory controller …”

    So, aside from no DDR2 support – which is inconsequential anyway since the G5 didn’t have it then either – you got almost everything wrong. Good job!

    I don’t know why I put down DDR2 but that’s not what I was thinking of (which is due in revision F), I had meant DDR500 and higher. The new memory controller allowed for support of much faster DDR memory without overclocking. Other than that, Revision E is exactly what I stated and more. Opterons/Athlons come in a variety of different cores. The article you linked to (and barefeats) was the Rev E “Troy” architecture published in March. AMD did not introduce Rev E “Venice” in their chips until summer.
    http://www.anandtech.com/cpuchipsets/showdoc.aspx?i=2469

    Can you answer me this question (without googling), what is the difference between an Athlon 4000 San Diego core and Venice core, or X2 4400 Manchestor core or Toledo core?

  46. I read it, AND IT’S WRONG. Here:
    “In general, the speedup in completion rate versus a single-cycle implementation that’s gained from pipelining is ideally equal to the number of pipeline stages. A four-stage pipeline yields a four-fold speedup in the completion rate versus single-cycle, a five-stage pipeline yields a five-fold speedup, a twelve-stage pipeline yields a twelve-fold speedup, and so on. This speedup is possible because the more pipeline stages there are in a processor, the more instructions the processor can work on simultaneously and the more instructions it can complete in a given period of time. So the more finely you can slice those four phases of the instruction’s lifecycle, the more of the hardware that’s used to implement those phases you can put to work at any given moment.”
    http://arstechnica.com/articles/paedia/cpu/pipelining-2.ars/1

    I urge you to read the rest of the article for yourself – even you might learn something.

    The problem with your link, is you don’t understand what was being discussed. From that very same article, I give you this quote:

    ….Note that the total execution time for each individual instruction is not changed by pipelining. It still takes an instruction 4ns to make it all the way through the processor; that 4ns can be split up into 4 clock cycles of 1ns each, or it can cover one longer clock cycle, but it’s still the same 4ns. Thus pipelining doesn’t speed up instruction execution time, but it does speed up program execution time (i.e. the number of nanoseconds that it takes to execute an entire program) by increasing the number of instructions finished per unit time.

    So you see, Odyssey, you’re still wrong….and you don’t even understand what was being discussed. Instructions per cycle are clock for clock SLOWER on the Pentium 4’s longer staged pipeline compared to their previous parts. Do you have a problem reading? I already posted this above already directly from Anandtech, but I’ll do it again:

    The 20-stage pipeline on the Pentium 4 is what allows it to hit higher clock speeds right off the bat without requiring a die shrink. It is for this reason that the Pentium 4 will debut at speeds of 1.4GHz and higher (we will talk more about clock speed in a bit). Before you let that number impress you too much, you have to realize that the 20-stage pipeline of the Pentium 4 also yields what is called a lower amount of Instructions Per Clock (IPC). A lower IPC basically means that you get less accomplished in a given amount of time when compared to a processor that has a higher IPC – pretty simple right?

    Well, there are a number of ways you make up for a lower IPC; one of the most obvious is to simply increase the clock speed, which Intel is definitely doing in this case. There isn’t a doubt that on any of the current benchmarks, if a 1GHz Pentium III were put up against a hypothetical 1GHz Pentium 4, the Pentium III would win because it can do more per clock than the Pentium 4.
    http://www.anandtech.com/showdoc.aspx?i=1301&p=3

    Highlights include Hyper Pipelined Technology, which enables the Pentium 4 processor to execute software instructions in a 20-stage pipeline, as compared to the 10-stage pipeline of the Pentium III processor. The 20-stage pipeline performs less work per clock cycle in exchange for higher clock speeds is the main reason the Pentium 4 has scaled this high so quickly. When built on the same manufacturing process, the Pentium 4 can scale to a higher clock speed much easier than its predecessor.
    http://www.hwextreme.com/reviews/processor/intel_2ghz/

  47. Here’s another good one:
    The number of stages (dot-units) modern processors break workloads into result in a certain number of things being done at a given time. And the way it

    works in silicon is like this: the smaller the workload per dot-unit, the faster it can be processed–because it’s doing less actual work per clock.

    Think of it like carrying water from a faucet to a small pond. In one instance you have 12 buckets that each hold a gallon of water. If you fill the buckets,

    move the buckets, and empty the buckets, you’ve moved 12 gallons of water in one trip. If the pond is 120 gallons, it will take 10 trips.

    For the P4 chip, you have either 20 or 31 buckets, but each is smaller, and you’ll be moving 20 or 31 buckets each trip. The net result is you have to fill

    more buckets and do more work each trip to get the same workload accomplished, since it takes more effort to move a bucket into place in front of the faucet,

    let it fill, then move it out of place, grab another bucket, put it in front of the faucet, and so on for 20 or 31 smaller buckets than it does to do the

    same with 12 larger buckets. It still takes 10 trips to fill the pond; you’re just doing more work each trip. But because you’re using less water in each

    bucket, each bucket is filled more quickly than the gallon-sized buckets.

    It boils down to a design tradeoff. Do you want to do less work at a faster pace, or more at a slower pace?

    Putting it together in reality
    Intel’s Pentium 4 and Xeon chips use a technology called Netburst, which uses either 20 or 31 pipeline stages. This means that a single thing, such as “Add A

    to B and store the results in A” is broken down into 20 or 31 separate stages. The same workload on any AMD64 chip (Athlon 64 or Opteron) takes 12 stages.

    This means that in order for an Intel processor to add A to B and store the result in A, it has to do 20 or 31 things. For AMD to complete the same workload,

    the processor only has to do 12 things.

    This is why an AMD chip operating at 2.8GHz processes data faster than an Intel chip running at 4.0GHz. Even though the Intel chip is going faster in MHz,

    it is doing less work per clock cycle.
    http://www.geek.com/news/geeknews/2005Dec/bch20051215033811.htm

    OK Odyssey, how many more damn links does your hard head require?

    Here’s what your own link says: “Modern day CPUs attempt to increase the efficiency of their pipelines by predicting what they will be asked to do next. This is a simplified explanation of the term Branch Tree Prediction. When a processor predicts correctly, everything goes according to plan but when an incorrect prediction is made, the processing cycle must start all over at the beginning of the pipeline. Because of this, a processor with a 10 stage pipeline has a lower penalty for a mis-predicted branch than that of a processor with a 20 stage pipeline. The longer the pipeline, the further back in the process you have to start over in order to make up for a mis-predicted branch.”

    Who’s having the problem with reading things again?

    You truly are ignorant Odyssey, and your level of ignorance truly astounds me. You are quoting a fact of branch prediction that is inherent to pipeline stages. Branch prediction misses are greater with longer pipelines….what exactly is Intel specific here? The same holds true with AMD’s 12 stage pipeline compared to the Pentium 3’s 10 stage. Is AMD’s branch prediction “lousy” as well?

    I think you have a serious issue with reading and understanding, Odyssey.

  48. What’s lousy about it is that it’s not up to the job of dealing with longer pipelines! I mean, you put the same technology in a shorter pipelined CPU (as is happening with Yonah) then the miss-hits don’t impose as much of a penalty. But, as per Intel practice for the last few years, they didn’t sweat the details when they decided to go with 20 staged Netburst. They did address the problem later, with the Prescott revisions, but in addition to beefing up the branch prediction they lengthened the pipeline AGAIN, to 30 stages. That kept it on pace with the older Northwood version, but AMD still beat it on Winstone tests (which are branchy in nature). Here:
    “In both tests, Prescott is essentially in a statistical dead heat with the 3.2GHz Northwood part — slightly behind in Business Winstone and slightly ahead in CC Winstone. Given the branchy nature of business applications, Prescott’s performance in Business Winstone is surprisingly good. Clearly the larger cache, improved branch prediction and better memory handling offsets the deeper pipeline.”

    It has nothing to do with “being up to the job of dealing with longer pipelines”. Your issue is not with branch prediction itself, but the number of pipeline stages…..you have absolutely no clue what envelopes the process of branch prediction, and here you are trying to pass it off as if you did. I ask you AGAIN, what SPECIFICALLY is lousy about Intel’s branch prediction algorithm? Can you even explain the method of how the algorithm processes each stage and what can be improved? Yeah, and don’t think about quoting articles your dumbass cannot comprehend, let me hear in your own words.

    What the hell’s a “smaller, longer stage pipeline” supposed to mean?? It’s one or the other genius.

    That just proved right there you dont have a damn clue what’s being discussed, you’re just trying to hard to mask your ignorance and hiding behind articles you do not understand (not to mention not support your arguments at all LOL). Otherwise you would have picked it up right away. “A longer staged pipeline will ALWAYS be less efficient than a shorter stage pipeline in terms of branch prediction, that is simple EE logic (10% missed prediction on an 11 stage CPU is less significant than a 20 stage CPU). Clock speed becomes less indirectly comparable just because of that fact (just as the article states, a 1Ghz Pentium 3 is faster than a 1Ghz Pentium 4). However, having a smaller, longer stage pipeline [increased number of smaller stages] allowed for much higher clockspeed to compensate. In contrast, an Opteron is a 12 stage pipelined CPU…the fact that it cannot (currently) hit beyond 3Ghz is not a coincidence.”

    Notice the bracket, did that help you a bit to understand? No? I’m not surprised one bit.

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.