Monday, February 8, 2016

More AMD Bristol Ridge SKUs leaked

Kristian Gocnik (@I_biT_MySeLf) tipped me off about new mobile Bristol Ridge SKUs, which appeared on usb.org as you can see [UPDATE: the entries have been removed now - visiting the pages may delete your only cached copy in the browser] here and here. That's the same site, where the first Bristol Ridge SKU (FX-9830P) appeared on. I put this together with information found in the leaked slide by Benchlife.info, which you can find in my blog post about a WEI result of an A10-9600P.

Table with leaked mobile Bristol Ridge OPNs

Using the mobile Carrizo SKUs, the leaked A10-9600P clock, and some sorting, it was easy to map the SKUs to the leaked slide's data. Kristian Gocnik tried it independently and we got the same mapping, except for a consumer A8-9500P he speculatively derived from the pro model, but which is missing on usb.org. So the resulting table likely represents what AMD is going to release as mobile Bristol Ridge chips for the FP4 socket later this year.

The model numbers likely simply jumped by one thousand from Carrizo's and an additional thirty points for the 35W variants. Carrizo's wide TDP ranges got split into 15W and 35W TDPs. This might help to avoid the confusion about 15W and 35W Carrizos laptops. The CPU base clocks jumped significantly, while CPU Turbo and (maximum) GPU clocks kind of matured with the fab process.

A reason for the jump has been given by AMD at ISSCC 2016, as EE Times reported:
"For its part, AMD engineers showed smart ways of squeezing as much as 15% more performance out of its Carrizo PC processor, simply by applying more aggressive power management to the 28nm design. The Bristol Ridge design was a study in using power management to overcome performance limits tied to heat, voltage and current."
Months after the first leaked WEI score, first true Bristol Ridge benchmarks will show, how this improvement translates into real world performance. Hopefully they get tested with dual channel memory, even if AMD or OEMs only provide single channel equipped/designed devices, as for the recent AnandTech Carrizo review.

BTW, there are lots of fresh Stoney Ridge Geekbench results in the Primate Labs' database.

Update: Of course, these are not OPNs, but SKUs. Added a warning as the linked usb.org entries are gone.

Monday, February 1, 2016

AMD Zeppelin CPU codename confirmed by patch and perhaps 32 cores per socket for Zen based MPUs, too

The Zeppelin codename, first mentioned on a leaked slide shown by Fudzilla, has been identified as a "family 17h model 00h" CPU by a patch on LKML.org. The interesting parts of the patch are:
AMD Zeppelin (Family 17h, Model 00h) introduces an instructionsretired performance counter which indicated byCPUID.8000_0008H:EBX[1]. And dedicated Instructions Retired register(MSR 0xC000_000E9) increments on once for every instruction retired.

There might even be a meaning behind the similarity of parts of the "Zen" and "Zeppelin" codenames.

An older patch on the same mailing list also gives a little more info about Zen:

On AMD Fam17h systems, the last level cache is not resident in Northbridge. Therefore, we cannot assign cpu_llc_id to same value as Node ID (as we have been doing currently)
We should rather look at the ApicID bits of the core to provide us the last level cache ID info. Doing that here.
The most interesting part describes the way, how the last level cache (LLC) ID is being calculated for Zen based MPUs:

+ core_complex_id = (apicid & ((1 << c->x86_coreid_bits) - 1)) >> 3;
+ per_cpu(cpu_llc_id, cpu) = (socket_id << 3) | core_complex_id;

"Core complex" should be similar to "compute unit" and has been used in some AMD patents already. The expression marked in red means a shift right by 3, which equals a division by 8. So with two logical cores per physical core due to SMT, a core complex should contain four Zen cores and a shared LLC.


The next line shows the socket ID being shifted left by 3, leaving 3 bits for the core complex ID, which suggests a maximum number of eight core complexes per socket, or 32 physical cores. This number should first be seen as a placeholder, but we've already seen rumours mentioning that many cores.

Tuesday, January 5, 2016

AMD A10-9600P: A Bristol Ridge laptop left some traces

One of my search strings got a hit today. Based on the already known model number 101 I found this Windows Experience Index CPU score of an AMD "A10-9600P" APU (Bristol Ridge). It's a quad core (2 modules) with 6 CUs and a reported base clock of 2.3GHz. A CPU score of 7.4 isn't that low, but doesn't tell us that much more, except that the system was able to finish the benchmark test. The result currently sits on page 7 of the filtered list linked by the table screenshot. You can also click the other images for the respective sources.

http://www.drivermax.com/driver/vista-rating/index_cpu.php?&start=7&filter_name=&filter_min_CpuScore=7.3&filter_max_CpuScore=7.5

 According to another page the listed model is a HP system, very likely a laptop.

http://www.drivermax.com/driver/laptop-rating/selvideo/Hewlett-Packard/InsydeH2O+EFI+BIOS/AMD+A10-9600P+RADEON+R5,+10+COMPUTE+CORES+4C%2B6G/pag0

As can be seen in the leaked slide below, the found CPU might be the a quad core with a cTDP of 12-15W.
Source: https://benchlife.info/amd-will-rename-excavator-to-bristol-ridge-12082015/

After the first listing of a Bristol Ridge SKU ("FX-9830P", thx @Onkel_Dithmeyer) and some BR ES traveling through the world, we now have a first Bristol Ridge APUs reporting the final product OPN instead of a typical ES string. Perhaps this is a system to be shown at CES?

Monday, November 23, 2015

AMD K12 looks to be at least a 4-wide design with SMT

An article about Zen and K12 by Yusuke Ohara gives a good overview of AMD's processor plans and a new and very interesting bit of information about AMD's high performance ARM design. As machine translators still struggle to provide clearly understandable translations of Japanese texts, multiple translators were tried and did not help. Therefore I asked the author to make sure that I got it right. He confirmed, that according to ARM officials, who are aware of the works of their architectural licensees, AMD is using at least a 4-wide design for their K12 core.



Jim Keller already said, that the smaller decoders for ARM instructions would leave room to add some performance improving features compared to x86. He also mentioned "a bigger engine" than in Zen. Looking at the microarchitecture diagram, one might ask, how AMD would utilize all these execution hardware, especially if there would be even more units and maybe even more than four instructions fetched and decoded per cycle. And given its target market, which is servers and datacenters, this might include one important feature: SMT. Some already speculated about that based on expectations, but there is AMD patent application US20150121046, which mentions SMT and its application in an AArch64 design very clearly and with many implementation details. This can be seen as an indicator of work being done for real products.

If K12 is a 4-wide or even wider SMT design similar to Zen (which is "only" 4-wide), this would put some substance behind Keller's announcements, which suggested many similarities between both designs. This is supported by the fact, that one of the inventors listed in the patents (Marius Evers) seemingly worked on both cores. Many other patents by him also cover both ARM and x86. He was also involved in one patent filed in 2007, which described a way to add SMT to the front end of a Bulldozer like module. SMT is not only useful to utilize execution units, if there are many of them. It also helps by keeping them busy, if there are multi-cycle FP instructions, branch mispredictions, or cache misses.

Of course, there are more differences between those two architectures than the ISAs alone, but many typical CPU components are either ISA-agnostic and reusable or could be adapted with much less effort than creating them from scratch. However, if it was done this way, such a strategy would not only have permitted AMD to make an efficient use of the limited R&D resources available, but it would have created a chance to produce a powerful ARM core for servers for an acceptable overhead. This is like applying SMT to R&D.

Wednesday, November 18, 2015

AMD Hierofalcon/Seattle shown at ARM TechCon

AMD presented some boards at ARM TechCon and thanks to ARMdevices.net there are two videos covering that stuff.

One video shows Red Hat's Jon Masters' explanation of AMD's Huskyboard, where (even if only printed on cardboard) you can have a nice closeup view of the chip (video screenshot):

AMD Seattle closeup

The second video shows real hardware at work, including SoftIron Overdrive 3000, and the Huskyboard in 3D:




Thursday, October 15, 2015

AMD Zen and K12 (ARM) tapeouts confirmed by LinkedIn profile

According to a LinkedIn profile, both Zen and K12 should have been taped out already. So this is a fact, as it isn't speculated based on sparse information. Interestingly the same guy (you have to find him yourself, if you need to), who only talks about CPU cores, mentions his working on 16nm and 14nm FinFet designs. So there will be one design made by TSMC and one by Globalfoundries. K12 by the first and Zen by the latter I suppose. And here is the snippet:


AMD's ARM-based "Hierofalcon" SoC sighted

On the same OSADL site, which once provided some first signs of life of a 2 GHz Jaguar based APU, there is now an engineering sample of one of AMD's ARM based embedded processors, called "Hierofalcon". The processor can be found in rack #a slot #3. According to the tables and logs, it also runs at 2 GHz and has 8 cores. If you haven't heard of that processor, just check these two slides. I actually included the first only because of the bird. ;)



I prepared some charts out of the numbers given there. If you check the site, you'll find no directly linked latency chart. But their 1337 page allows to compare rack #a slot #3 to some other rack/slot and returns the missed latency chart of the "Hierofalcon", which looks like this (updated daily):
Latency plot of AMD "Hierofalcon" ES
Latency plot of AMD "Hierofalcon" ES
The only available performance numbers are some daily updated Unixbench results. So I took them, combined rack names with CPU strings, sorted them, choose some CPUs for comparison, and normalized the results to the CPU in question.


The first chart already shows, that on a per clock basis AMD's other CPUs already lag behind in simple integer code of the old kDhrystone benchmark. The floating point based Whetstone benchmark draws a somewhat different picture with more equally distributed per clock performances except that of the old K10 based Phenom II. The next three benchmarks Execl (not Excel!), kCopy, and kPipe test OS functions like spawning processes, doing file copying or using the pipe. The Index is a combined result.

In the next chart we can see the raw performance of all cores, only normalized again to Hierofalcon.


Even then the 8 cores of the ARM based processor have a good standing in the first two benchmarks, while in the OS benchmark, it roughly keeps up with Kaveri and Bulldozer, both running at much higher clock speeds.

The ARM based CPUs are meant to put many lower power cores together. To have a first impression of that effect, I used the given TDP numbers as the only metric available for all CPUs. Here are the power efficiency numbers:



I think in this case, the Hierofalcon bars are really easy to spot, even though I used the max listed TDP of 30W. Only the already power optimized Sandy Bridge variants and the 9W Kabini are able to keep up in some of the tests. And of course, real power measurements would shift the numbers a bit.

Of course, many (including me) would like to see more interesting benchmarks, but these are the first numbers we've got and they aren't bad at all.