C99 inline without static does not emit a function body, so aarch64
builds at -O0/-Os fail to link DetectArmFeatures (undefined reference
to CPU_QueryAES / CPU_QuerySHA2). Use VC_INLINE like the other helpers
in cpu.c.
Signed-off-by: Ville Takio <ville+git@takio.fi>
With NOASM=1 on x86/x64 the assembler module Aes_hw_cpu is not built,
but cpu.h still defines TC_AES_HW_CPU, so Cipher.cpp references
aes_hw_cpu_encrypt/decrypt. cpu.c also still declares and calls
TrySHA256, which Sha2Intel.c omits when CRYPTOPP_DISABLE_ASM is set
(NOASM passes CRYPTOPP_DISABLE_X86ASM, which implies it), and this
happens whenever __SHA__ is defined (-msha, -march=native) or
CRYPTOPP_SHANI_AVAILABLE is set.
- cpu.h: define TC_AES_HW_CPU on x86 only without CRYPTOPP_DISABLE_ASM
- cpu.c: declare and call TrySHA256 only under the same condition
Sha2Intel.c uses to build it (not _UEFI, not CRYPTOPP_DISABLE_ASM)
Both follow the compiler target, so cross builds need no ARCH override.
Signed-off-by: Ville Takio <ville+git@takio.fi>
BMI2 support is advertised by CPUID leaf 7, subleaf 0, EBX bit 8. The previous early assignment used CPUID leaf 1 EBX bit 8, which is not the BMI2 feature bit and could leave a bogus fallback value before vendor-specific leaf 7 detection.
Keep BMI2 detection based on the leaf 7 result only. Unlike AVX2, BMI2 is GPR-only and does not require an OS/XCR0 state gate.
Also save the max basic CPUID leaf immediately after CPUID leaf 0. The AMD/Hygon path reuses the cpuid buffer for leaf 0x80000005 before checking whether leaf 7 is available, so using the saved max basic leaf prevents RDSEED, AVX2, and BMI2 detection from being skipped because that buffer was clobbered.
AVX2 support is advertised by CPUID leaf 7, subleaf 0, EBX bit 5. The previous early assignment used cpuid1[1] bit 5, which is CPUID leaf 1 EBX and is not the AVX2 feature bit.
Record the leaf 7 AVX2 bit separately and assign g_hasAVX2 only after vendor-specific detection has completed. The final value is now gated by g_hasAVX, which reflects the OS/XCR0 AVX state check, so AVX2 code is not selected unless both the CPU and OS state support it.