A 49-bit prime chosen to fit the instruction
MamaBearZKP picks its field to match AVX-512IFMA and reports up to 18× over Plonky3 on a single thread.
2 minZero-Knowledge Proofs
Jipeng Zhang, Yanpei Guo, Tao Lu, Hao Cheng and Jiaheng Zhang, of the National University of Singapore and Shandong University, revised MamaBearZKP on 19 August. The paper co-designs the prime field and the proving stack together rather than treating them as separate layers.
The field is p = 2^49 − 2^34 + 1, and the choice is not arbitrary: 49 bits is what fits AVX-512IFMA's integer multiply-accumulate path on modern processors. The optimisation target is the instruction, and the mathematics is selected to suit it.
The numbers, with their baseline
Against a Goldilocks baseline the authors report single-thread speedups of up to 42×, 33×, 15×, 21× and 21× across different proving operations, rising to 64×, 47×, 81×, 45× and 45× on eight threads. Against Plonky3 the figure is up to 18× on a single thread.
The Plonky3 comparison is the one that carries weight. Goldilocks is a reference point the authors chose; Plonky3 is a system people actually run, and an 18× single-thread gap against it is the claim a reader can go and test.
The caveat is in the premise. A field tuned to one instruction set is fast where that instruction set exists. On hardware without AVX-512IFMA — including much of what runs in cloud instances by default — the argument has to be made again from scratch.
Retold from IACR ePrint. This is a summary in our own words; follow the link for the original reporting.