crypto: fix ARMv8 slow-hash inline assembly constraints
What changed, and why it matters
This commit fixes the way a Monero cryptographic function for ARMv8 processors describes its inline assembly code to the C compiler. Previously, the assembly block did not properly tell the compiler which memory and CPU registers it reads and writes. This can lead to subtle bugs where the compiler optimizes code incorrectly around the assembly, potentially causing wrong hash results or memory corruption on ARMv8 devices. The fix adds proper input/output constraints and a clobber list so the compiler knows exactly what the assembly touches.
Treat as a low-to-moderate reliability/correctness fix. ARMv8 Monero nodes/miners should update to avoid potential incorrect hashing or crashes from compiler misoptimization. No immediate emergency response is warranted absent evidence of exploitable memory corruption.
Security signals we found
Inline assembly missing proper input/output/memory constraints
Hardcoded register usage (w2) without compiler coordination
Potential for compiler misoptimization or register/memory corruption
ARMv8-specific cryptographic code path affected
No explicit security disclosure or CVE referenced in commit
Evidence from the diff
The patch updates the ARMv8 AES key expansion inline assembly in src/crypto/slow-hash.c. The old code used an unannotated asm block with hardcoded register operands and no output operands, clobbers, or memory constraints. It also used a hardcoded w2 register for the loop counter without declaring it. The new version introduces local pointer variables (key_ptr, expanded_key_ptr, rcon_ptr) and a count variable, binds them to operands with proper early-clobber and read/write constraints (+&r, =&r), declares memory inputs/outputs for rcon, key, and expandedKey, and lists the SIMD registers and condition codes in the clobber list. This is a correctness fix for compiler-assembly interaction, not a change to the cryptographic algorithm itself.
Changed components
src/crypto/slow-hash.cARMv8 slow-hash implementationAES key expansion inline assemblyInspect captured patch +16 / −3
diff --git a/src/crypto/slow-hash.c b/src/crypto/slow-hash.c
index 305786a..9aa6e43 100644
--- a/src/crypto/slow-hash.c
+++ b/src/crypto/slow-hash.c
@@ -1166,12 +1166,17 @@ static const int rcon[] = {
0x01,0x01,0x01,0x01,
0x0c0f0e0d,0x0c0f0e0d,0x0c0f0e0d,0x0c0f0e0d, // rotate-n-splat
0x1b,0x1b,0x1b,0x1b };
+const uint8_t *key_ptr = key;
+uint8_t *expanded_key_ptr = expandedKey;
+const int *rcon_ptr = rcon;
+int count;
+
__asm__(
" eor v0.16b,v0.16b,v0.16b\n"
" ld1 {v3.16b},[%0],#16\n"
" ld1 {v1.4s,v2.4s},[%2],#32\n"
" ld1 {v4.16b},[%0]\n"
-" mov w2,#5\n"
+" mov %w3,#5\n"
" st1 {v3.4s},[%1],#16\n"
"\n"
"1:\n"
@@ -1179,7 +1184,7 @@ __asm__(
" ext v5.16b,v0.16b,v3.16b,#12\n"
" st1 {v4.4s},[%1],#16\n"
" aese v6.16b,v0.16b\n"
-" subs w2,w2,#1\n"
+" subs %w3,%w3,#1\n"
"\n"
" eor v3.16b,v3.16b,v5.16b\n"
" ext v5.16b,v0.16b,v5.16b,#12\n"
@@ -1205,7 +1210,15 @@ __asm__(
" eor v4.16b,v4.16b,v6.16b\n"
" b 1b\n"
"\n"
-"2:\n" : : "r"(key), "r"(expandedKey), "r"(rcon));
+"2:\n"
+ : "+&r"(key_ptr),
+ "+&r"(expanded_key_ptr),
+ "+&r"(rcon_ptr),
+ "=&r"(count),
+ "=m"(*(uint8_t (*)[176]) expandedKey)
+ : "m"(*(const int (*)[8]) rcon),
+ "m"(*(const uint8_t (*)[32]) key)
+ : "v0", "v1", "v2", "v3", "v4", "v5", "v6", "cc");
}
/* An ordinary AES round is a sequence of SubBytes, ShiftRows, MixColumns, AddRoundKey. There
Why this scored 31/100
Community notes
Notes can correct, qualify, or add evidence to the AI analysis. Every note shown here has been validated by a human moderator.
The AI analysis stands alone for now. Submit a note if you can add evidence or important context.