InstLatX64

@instlatx64.bsky.social

x86/x64, SIMD, #AVX512, "Aha!" moments. I have been writing code since 1986. Budapest, Europe https://instlatx64.github.io/InstLatx64/

#AMD [1]: - #Zen8 EPYC: Ravenna - #Zen7 EPYC: -- for AI host Ferrara -- for Agentic Sandbox Fidenza - #Zen6 [2], [3]: -- desktop #OlympicRidge B80F80 #AM5 -- mobile #Medusa1 B80F00 AM5/FP10 -- mobile #Medusa2 BE0F00 #FP10 #Intel - #GriffinCove for RZL [4] - #CopperShark [5] - #ThunderHawk [6] 1/3

InstLatX64@instlatx64.bsky.social · last wk.

InstLatX64 refresh: #Intel: - 2C+4c #Core 7 350, Core 5 320 #WildcatLake D0651 CPUID dump GitHub: github.com/InstLatx64/I...

It's really hard to let go of this 3x11 #Zen6c #Monarch CCD configuration idea, because it simply explains the unusual number of cores in CCDs by disabling rows and columns. Okay, that's not great for L3 latency, but I think something similar is needed for the risk of the brand new TSMC N2 process.

BildBild

Unfortunately, the Perfmon counters of the #Intel #NovaLake INT cluster inherited the grouped nature from the older *Coves with a central scheduler, so the number of ports used here is unknown (except for the 3 JEUs).

BildBild
InstLatX64@instlatx64.bsky.social · 3w ago

#Intel perf tool got #NovaLake CPUID 300F10 300F30 support: According to this commit, while the #ArcticWolf config is quite similar to #Skymont / #Darkmont (8 INT+4 FP, FP probably 128b->256b), the #CoyoteCove FP has increased to 6 ports from 4 in #LionCove / #CougarCove. github.com/intel/perfmo...

#AMD refreshed the "AMD64 Architecture Programmer's Manual, Volumes 1-5" 40332 to 4.10, all-in-one: Vol1 24592-Rev. 3.25-May 2026 Vol2 24593-Rev. 3.45-Jul 2026 Vol3 24594-Rev. 3.38-Jul 2026 Vol4 26568-Rev. 3.27-Jul 2026 Vol5 26569-Rev. 3.16-Nov 2021 docs.amd.com/v/u/en-US/40...

BildBildBildBild
InstLatX64@instlatx64.bsky.social · 5mo ago

#AMD refreshed the "AMD64 Architecture Programmer's Manual, Volumes 1-5" 40332 to 4.09, all-in-one: Vol1 24592-Rev. 3.24-Aug 2025 Vol2 24593-Rev. 3.44-Mar 2026 Vol3 24594-Rev. 3.37-Jul 2025 Vol4 26568-Rev. 3.26-Jan 2026 Vol5 26569-Rev. 3.16-Nov 2021 docs.amd.com/v/u/en-US/40...

Until the 84-core #AMD #Sorano #EPYC 8635P came to mind (12 CCD x 7-core), this 176-core (8 CCD x 22-core) #Zen6c seemed completely incomprehensible without a 3rd CCD type. But supposing 11x3 core within the #Monarch CCD makes this a yield optimization. blogs.microsoft.com/blog/2026/07...

BildBild
InstLatX64@instlatx64.bsky.social · 3mo ago

This SPEC pdf mentions two #AMD #Sorano #EPYC 8xx5 SKUs: AMD EPYC 8635P AMD EPYC 8225P www.spec.org/sert2/ISO_Re... 8635P seems the 84-core TOP SKU: www.connection.com/product/hpe-...

CPUID.1Ah.EAX: 20000000: #Tremont 20000001: #Gracemont 20000002: #Crestmont 20000003: #Skymont 20000004: #Darkmont 20000005: #ArcticWolf 40000000: #SunnyCove 40000001: #GoldenCove 40000002: #RedwoodCove 40000003: #LionCove 40000004: #CougarCove 40000005: #CoyoteCove GitHub github.com/intel/perfmo...

Bild
InstLatX64@instlatx64.bsky.social · last yr.

#Intel perfmon updated with #PantherLake CPUID.1Ah.EAX values: 20000000 #Tremont 20000001 #Gracemont 20000002 #Crestmont 20000003 #Skymont 20000004 #Darkmont 40000000 #SunnyCove 40000001 #GoldenCove 40000002 #RedwoodCove 40000003 #LionCove 40000004 #CougarCove github.com/intel/perfmo...

#Intel perf tool got #NovaLake CPUID 300F10 300F30 support: According to this commit, while the #ArcticWolf config is quite similar to #Skymont / #Darkmont (8 INT+4 FP, FP probably 128b->256b), the #CoyoteCove FP has increased to 6 ports from 4 in #LionCove / #CougarCove. github.com/intel/perfmo...

BildBild
InstLatX64@instlatx64.bsky.social · 12mo ago

#Intel #PantherLake got perfmon support: github.com/intel/perfmo... According to this #CougarCove and #Darkmont is similar to #LionCove and #Skymont, just the front-end/branch prediction is different.

Another interesting GCC patch: disabling the memory-form of NDD #Intel #APX instructions. #DiamondRapids is not mentioned, perhaps this is an #ArcticWolf related issue? i386: Add tuning to disable memory-form NDD gcc.gnu.org/git/?p=gcc.g...

BildBild
InstLatX64@instlatx64.bsky.social · last mo.

An interesting #APX GCC patch: i386: Disable SETcc.ZU generation on DMR/NVL via tune flag gcc.gnu.org/git/?p=gcc.g... Strange, isn't a 6 byte EVEX better than a 2/3/4 byte XOR + 3/4 byte SETcc? #DiamondRapids #NovaLake #PantherCove

#Intel released the 62nd edition of the ISA Extensions Reference with. #UTMR and #AMX_TF32 support removed (it means only one instruction: TMMULTF32PS) Download: cdrdv2-public.intel.com/922690/31943... #DiamondRapids #NovaLake #PantherCove #CoyoteCove #ArcticWolf

Bild
InstLatX64@instlatx64.bsky.social · 5mo ago

#Intel released the 61st edition of the ISA Extensions Reference with PerfMon Masking and clarifications. Download: cdrdv2-public.intel.com/915637/31943... #DiamondRapids #NovaLake #WildcatLake #PantherCove #CoyoteCove #ArcticWolf CPUID.07h.01.EDX[22]= #SEC_TEE_ATTESTATION

Strange. #AMD's documentation page lists both #SP7 and #SP8 's embedded #Zen6 #EPYC as 9006 instead of 11006/10006 (or B006/A006 if they've switched to the more practical hex notation). docs.amd.com/search/all?v...

Bild
InstLatX64@instlatx64.bsky.social · last yr.

New #Intel CPUIDs [1]: - #NovaLake 300F10 (Family 18 model 1) - #NovaLakeL 300F30 (Family 18 model 3) New dump: -Core Ultra 9 275HX C0662 #ArrowLakeHX New #AMD #Zen6 Venice #EPYC socket assignments [2]: -B50F00, BC0F00 #SP7 16 m. c. -B90F00, BA0F00 #SP8 8 m. c. GitHub: github.com/InstLatx64/I... 1/2