Skip to content

feat: napi rebuild - #6660

Closed
matthewkeil wants to merge 31 commits into
unstablefrom
mkeil/napi-rebuild-2-rebase-unstable-final-bls
Closed

feat: napi rebuild#6660
matthewkeil wants to merge 31 commits into
unstablefrom
mkeil/napi-rebuild-2-rebase-unstable-final-bls

Conversation

@matthewkeil

Copy link
Copy Markdown
Member

Description

Uses the latest published versions of chainsafe/bls@8.0.0 and chainsafe/blst@1.0.0.

Includes PR comments moved from #6616

Base of PR was branch mkeil/test-blst-7 from #6362 and has the added verification randomness for same messages added. Retains workers and herumi still.

Currently deployed to feat1 for metrics collection

@codecov

codecov Bot commented Apr 11, 2024

Copy link
Copy Markdown

Codecov Report

Merging #6660 (9d82670) into unstable (5ccae1c) will increase coverage by 0.18%.
Report is 5 commits behind head on unstable.
The diff coverage is 95.65%.

❗ Current head 9d82670 differs from pull request most recent head fc3ce71. Consider uploading reports for the commit fc3ce71 to get more accurate results

Additional details and impacted files
@@             Coverage Diff              @@
##           unstable    #6660      +/-   ##
============================================
+ Coverage     61.64%   61.82%   +0.18%     
============================================
  Files           556      556              
  Lines         58782    59073     +291     
  Branches       1886     1898      +12     
============================================
+ Hits          36236    36524     +288     
- Misses        22505    22506       +1     
- Partials         41       43       +2     

@github-actions

github-actions Bot commented Apr 11, 2024

Copy link
Copy Markdown
Contributor

Performance Report

✔️ no performance regression detected

Full benchmark results
Benchmark suite Current: ebe50f7 Previous: 2dae605 Ratio
getPubkeys - index2pubkey - req 1000 vs - 250000 vc 1.2045 ms/op 1.0150 ms/op 1.19
getPubkeys - validatorsArr - req 1000 vs - 250000 vc 76.673 us/op 96.065 us/op 0.80
BLS verify - blst-native 1.2199 ms/op 1.5204 ms/op 0.80
BLS verifyMultipleSignatures 3 - blst-native 2.3315 ms/op 3.1984 ms/op 0.73
BLS verifyMultipleSignatures 8 - blst-native 5.0031 ms/op 6.9859 ms/op 0.72
BLS verifyMultipleSignatures 32 - blst-native 18.091 ms/op 26.728 ms/op 0.68
BLS verifyMultipleSignatures 64 - blst-native 35.555 ms/op 51.190 ms/op 0.69
BLS verifyMultipleSignatures 128 - blst-native 70.463 ms/op 104.13 ms/op 0.68
BLS deserializing 10000 signatures 824.18 ms/op 1.0891 s/op 0.76
BLS deserializing 100000 signatures 8.2507 s/op 10.512 s/op 0.78
BLS verifyMultipleSignatures - same message - 3 - blst-native 1.2395 ms/op 1.6456 ms/op 0.75
BLS verifyMultipleSignatures - same message - 8 - blst-native 1.3953 ms/op 1.7367 ms/op 0.80
BLS verifyMultipleSignatures - same message - 32 - blst-native 2.1544 ms/op 2.6032 ms/op 0.83
BLS verifyMultipleSignatures - same message - 64 - blst-native 3.1646 ms/op 3.9652 ms/op 0.80
BLS verifyMultipleSignatures - same message - 128 - blst-native 5.1930 ms/op 6.3191 ms/op 0.82
BLS aggregatePubkeys 32 - blst-native 26.450 us/op 27.822 us/op 0.95
BLS aggregatePubkeys 128 - blst-native 102.53 us/op 117.11 us/op 0.88
notSeenSlots=1 numMissedVotes=1 numBadVotes=10 58.919 ms/op 69.618 ms/op 0.85
notSeenSlots=1 numMissedVotes=0 numBadVotes=4 64.053 ms/op 65.155 ms/op 0.98
notSeenSlots=2 numMissedVotes=1 numBadVotes=10 36.838 ms/op 58.821 ms/op 0.63
getSlashingsAndExits - default max 149.06 us/op 174.36 us/op 0.85
getSlashingsAndExits - 2k 438.79 us/op 442.95 us/op 0.99
proposeBlockBody type=full, size=empty 6.0594 ms/op 8.5465 ms/op 0.71
isKnown best case - 1 super set check 302.00 ns/op 507.00 ns/op 0.60
isKnown normal case - 2 super set checks 294.00 ns/op 494.00 ns/op 0.60
isKnown worse case - 16 super set checks 285.00 ns/op 472.00 ns/op 0.60
InMemoryCheckpointStateCache - add get delete 5.9800 us/op 6.1220 us/op 0.98
validate api signedAggregateAndProof - struct 2.4914 ms/op 2.9111 ms/op 0.86
validate gossip signedAggregateAndProof - struct 2.4896 ms/op 3.2857 ms/op 0.76
validate gossip attestation - vc 640000 1.3320 ms/op 1.6321 ms/op 0.82
batch validate gossip attestation - vc 640000 - chunk 32 159.05 us/op 181.01 us/op 0.88
batch validate gossip attestation - vc 640000 - chunk 64 140.09 us/op 166.35 us/op 0.84
batch validate gossip attestation - vc 640000 - chunk 128 132.92 us/op 159.27 us/op 0.83
batch validate gossip attestation - vc 640000 - chunk 256 132.91 us/op 175.93 us/op 0.76
pickEth1Vote - no votes 1.2175 ms/op 1.4227 ms/op 0.86
pickEth1Vote - max votes 10.129 ms/op 10.550 ms/op 0.96
pickEth1Vote - Eth1Data hashTreeRoot value x2048 20.398 ms/op 22.341 ms/op 0.91
pickEth1Vote - Eth1Data hashTreeRoot tree x2048 27.538 ms/op 34.028 ms/op 0.81
pickEth1Vote - Eth1Data fastSerialize value x2048 626.32 us/op 903.90 us/op 0.69
pickEth1Vote - Eth1Data fastSerialize tree x2048 6.3090 ms/op 5.9811 ms/op 1.05
bytes32 toHexString 494.00 ns/op 754.00 ns/op 0.66
bytes32 Buffer.toString(hex) 286.00 ns/op 400.00 ns/op 0.71
bytes32 Buffer.toString(hex) from Uint8Array 411.00 ns/op 613.00 ns/op 0.67
bytes32 Buffer.toString(hex) + 0x 280.00 ns/op 359.00 ns/op 0.78
Object access 1 prop 0.15600 ns/op 0.24700 ns/op 0.63
Map access 1 prop 0.14300 ns/op 0.17800 ns/op 0.80
Object get x1000 7.3530 ns/op 8.5440 ns/op 0.86
Map get x1000 0.71400 ns/op 0.94800 ns/op 0.75
Object set x1000 48.464 ns/op 63.211 ns/op 0.77
Map set x1000 38.508 ns/op 48.132 ns/op 0.80
Return object 10000 times 0.22910 ns/op 0.27110 ns/op 0.85
Throw Error 10000 times 3.7301 us/op 4.7553 us/op 0.78
fastMsgIdFn sha256 / 200 bytes 3.1770 us/op 5.6090 us/op 0.57
fastMsgIdFn h32 xxhash / 200 bytes 271.00 ns/op 531.00 ns/op 0.51
fastMsgIdFn h64 xxhash / 200 bytes 332.00 ns/op 606.00 ns/op 0.55
fastMsgIdFn sha256 / 1000 bytes 11.016 us/op 21.365 us/op 0.52
fastMsgIdFn h32 xxhash / 1000 bytes 398.00 ns/op 815.00 ns/op 0.49
fastMsgIdFn h64 xxhash / 1000 bytes 407.00 ns/op 1.0990 us/op 0.37
fastMsgIdFn sha256 / 10000 bytes 100.38 us/op 165.69 us/op 0.61
fastMsgIdFn h32 xxhash / 10000 bytes 1.8710 us/op 3.1860 us/op 0.59
fastMsgIdFn h64 xxhash / 10000 bytes 1.2930 us/op 1.7150 us/op 0.75
send data - 1000 256B messages 19.849 ms/op 26.598 ms/op 0.75
send data - 1000 512B messages 24.093 ms/op 56.860 ms/op 0.42
send data - 1000 1024B messages 40.365 ms/op 78.640 ms/op 0.51
send data - 1000 1200B messages 39.326 ms/op 50.884 ms/op 0.77
send data - 1000 2048B messages 49.476 ms/op 66.506 ms/op 0.74
send data - 1000 4096B messages 44.139 ms/op 57.605 ms/op 0.77
send data - 1000 16384B messages 112.41 ms/op 142.12 ms/op 0.79
send data - 1000 65536B messages 467.79 ms/op 514.81 ms/op 0.91
enrSubnets - fastDeserialize 64 bits 1.2330 us/op 1.6880 us/op 0.73
enrSubnets - ssz BitVector 64 bits 417.00 ns/op 645.00 ns/op 0.65
enrSubnets - fastDeserialize 4 bits 181.00 ns/op 262.00 ns/op 0.69
enrSubnets - ssz BitVector 4 bits 442.00 ns/op 610.00 ns/op 0.72
prioritizePeers score -10:0 att 32-0.1 sync 2-0 214.59 us/op 263.96 us/op 0.81
prioritizePeers score 0:0 att 32-0.25 sync 2-0.25 275.72 us/op 414.16 us/op 0.67
prioritizePeers score 0:0 att 32-0.5 sync 2-0.5 371.62 us/op 500.15 us/op 0.74
prioritizePeers score 0:0 att 64-0.75 sync 4-0.75 599.52 us/op 831.29 us/op 0.72
prioritizePeers score 0:0 att 64-1 sync 4-1 720.16 us/op 985.31 us/op 0.73
array of 16000 items push then shift 1.6551 us/op 2.0290 us/op 0.82
LinkedList of 16000 items push then shift 9.3070 ns/op 10.853 ns/op 0.86
array of 16000 items push then pop 103.25 ns/op 123.11 ns/op 0.84
LinkedList of 16000 items push then pop 8.9440 ns/op 10.493 ns/op 0.85
array of 24000 items push then shift 2.4642 us/op 3.3183 us/op 0.74
LinkedList of 24000 items push then shift 9.1590 ns/op 12.620 ns/op 0.73
array of 24000 items push then pop 148.43 ns/op 174.40 ns/op 0.85
LinkedList of 24000 items push then pop 9.0860 ns/op 11.093 ns/op 0.82
intersect bitArray bitLen 8 5.8470 ns/op 7.2830 ns/op 0.80
intersect array and set length 8 67.561 ns/op 88.935 ns/op 0.76
intersect bitArray bitLen 128 35.599 ns/op 43.968 ns/op 0.81
intersect array and set length 128 977.31 ns/op 1.3205 us/op 0.74
bitArray.getTrueBitIndexes() bitLen 128 1.7370 us/op 2.3680 us/op 0.73
bitArray.getTrueBitIndexes() bitLen 248 2.9260 us/op 3.6410 us/op 0.80
bitArray.getTrueBitIndexes() bitLen 512 6.5780 us/op 7.1860 us/op 0.92
Buffer.concat 32 items 1.0760 us/op 1.3940 us/op 0.77
Uint8Array.set 32 items 1.8930 us/op 2.5700 us/op 0.74
Set add up to 64 items then delete first 4.9825 us/op 6.0638 us/op 0.82
OrderedSet add up to 64 items then delete first 6.5572 us/op 8.5510 us/op 0.77
Set add up to 64 items then delete last 5.3805 us/op 6.0770 us/op 0.89
OrderedSet add up to 64 items then delete last 6.9925 us/op 7.8513 us/op 0.89
Set add up to 64 items then delete middle 5.3931 us/op 5.7263 us/op 0.94
OrderedSet add up to 64 items then delete middle 8.4389 us/op 9.3414 us/op 0.90
Set add up to 128 items then delete first 10.778 us/op 11.726 us/op 0.92
OrderedSet add up to 128 items then delete first 14.210 us/op 16.755 us/op 0.85
Set add up to 128 items then delete last 10.651 us/op 12.055 us/op 0.88
OrderedSet add up to 128 items then delete last 12.867 us/op 16.737 us/op 0.77
Set add up to 128 items then delete middle 10.421 us/op 11.795 us/op 0.88
OrderedSet add up to 128 items then delete middle 19.257 us/op 21.980 us/op 0.88
Set add up to 256 items then delete first 20.892 us/op 24.904 us/op 0.84
OrderedSet add up to 256 items then delete first 30.401 us/op 33.159 us/op 0.92
Set add up to 256 items then delete last 20.669 us/op 26.453 us/op 0.78
OrderedSet add up to 256 items then delete last 25.273 us/op 33.354 us/op 0.76
Set add up to 256 items then delete middle 21.040 us/op 24.836 us/op 0.85
OrderedSet add up to 256 items then delete middle 51.715 us/op 60.362 us/op 0.86
transfer serialized Status (84 B) 1.9910 us/op 2.0340 us/op 0.98
copy serialized Status (84 B) 1.4570 us/op 1.6180 us/op 0.90
transfer serialized SignedVoluntaryExit (112 B) 2.1090 us/op 2.4160 us/op 0.87
copy serialized SignedVoluntaryExit (112 B) 1.4990 us/op 1.7420 us/op 0.86
transfer serialized ProposerSlashing (416 B) 2.2320 us/op 4.0120 us/op 0.56
copy serialized ProposerSlashing (416 B) 1.9590 us/op 3.4060 us/op 0.58
transfer serialized Attestation (485 B) 2.3000 us/op 3.5300 us/op 0.65
copy serialized Attestation (485 B) 2.0240 us/op 2.9710 us/op 0.68
transfer serialized AttesterSlashing (33232 B) 4.2460 us/op 3.0150 us/op 1.41
copy serialized AttesterSlashing (33232 B) 20.169 us/op 10.062 us/op 2.00
transfer serialized Small SignedBeaconBlock (128000 B) 3.8410 us/op 3.5420 us/op 1.08
copy serialized Small SignedBeaconBlock (128000 B) 18.913 us/op 27.523 us/op 0.69
transfer serialized Avg SignedBeaconBlock (200000 B) 4.0070 us/op 3.7110 us/op 1.08
copy serialized Avg SignedBeaconBlock (200000 B) 26.130 us/op 32.143 us/op 0.81
transfer serialized BlobsSidecar (524380 B) 4.0850 us/op 4.1660 us/op 0.98
copy serialized BlobsSidecar (524380 B) 87.549 us/op 104.21 us/op 0.84
transfer serialized Big SignedBeaconBlock (1000000 B) 3.8720 us/op 3.9180 us/op 0.99
copy serialized Big SignedBeaconBlock (1000000 B) 179.28 us/op 185.93 us/op 0.96
pass gossip attestations to forkchoice per slot 4.8419 ms/op 5.7482 ms/op 0.84
forkChoice updateHead vc 100000 bc 64 eq 0 697.29 us/op 1.0062 ms/op 0.69
forkChoice updateHead vc 600000 bc 64 eq 0 4.3991 ms/op 5.2902 ms/op 0.83
forkChoice updateHead vc 1000000 bc 64 eq 0 7.2501 ms/op 8.9346 ms/op 0.81
forkChoice updateHead vc 600000 bc 320 eq 0 4.2495 ms/op 6.6889 ms/op 0.64
forkChoice updateHead vc 600000 bc 1200 eq 0 4.4542 ms/op 5.7342 ms/op 0.78
forkChoice updateHead vc 600000 bc 7200 eq 0 5.4163 ms/op 7.9181 ms/op 0.68
forkChoice updateHead vc 600000 bc 64 eq 1000 11.080 ms/op 16.578 ms/op 0.67
forkChoice updateHead vc 600000 bc 64 eq 10000 12.000 ms/op 15.482 ms/op 0.78
forkChoice updateHead vc 600000 bc 64 eq 300000 15.731 ms/op 19.611 ms/op 0.80
computeDeltas 500000 validators 300 proto nodes 6.7018 ms/op 7.4105 ms/op 0.90
computeDeltas 500000 validators 1200 proto nodes 6.5881 ms/op 7.4659 ms/op 0.88
computeDeltas 500000 validators 7200 proto nodes 6.4470 ms/op 7.4208 ms/op 0.87
computeDeltas 750000 validators 300 proto nodes 9.7276 ms/op 11.870 ms/op 0.82
computeDeltas 750000 validators 1200 proto nodes 9.6379 ms/op 10.689 ms/op 0.90
computeDeltas 750000 validators 7200 proto nodes 9.6318 ms/op 10.646 ms/op 0.90
computeDeltas 1400000 validators 300 proto nodes 18.770 ms/op 21.929 ms/op 0.86
computeDeltas 1400000 validators 1200 proto nodes 18.470 ms/op 22.379 ms/op 0.83
computeDeltas 1400000 validators 7200 proto nodes 18.257 ms/op 22.564 ms/op 0.81
computeDeltas 2100000 validators 300 proto nodes 27.538 ms/op 34.405 ms/op 0.80
computeDeltas 2100000 validators 1200 proto nodes 27.750 ms/op 33.082 ms/op 0.84
computeDeltas 2100000 validators 7200 proto nodes 27.160 ms/op 32.476 ms/op 0.84
altair processAttestation - 250000 vs - 7PWei normalcase 2.0911 ms/op 2.5186 ms/op 0.83
altair processAttestation - 250000 vs - 7PWei worstcase 3.5258 ms/op 3.4951 ms/op 1.01
altair processAttestation - setStatus - 1/6 committees join 168.59 us/op 227.82 us/op 0.74
altair processAttestation - setStatus - 1/3 committees join 339.39 us/op 405.16 us/op 0.84
altair processAttestation - setStatus - 1/2 committees join 459.80 us/op 551.76 us/op 0.83
altair processAttestation - setStatus - 2/3 committees join 549.70 us/op 648.62 us/op 0.85
altair processAttestation - setStatus - 4/5 committees join 857.45 us/op 833.23 us/op 1.03
altair processAttestation - setStatus - 100% committees join 907.64 us/op 1.0023 ms/op 0.91
altair processBlock - 250000 vs - 7PWei normalcase 8.0035 ms/op 8.4441 ms/op 0.95
altair processBlock - 250000 vs - 7PWei normalcase hashState 34.666 ms/op 35.832 ms/op 0.97
altair processBlock - 250000 vs - 7PWei worstcase 45.529 ms/op 44.437 ms/op 1.02
altair processBlock - 250000 vs - 7PWei worstcase hashState 103.76 ms/op 104.12 ms/op 1.00
phase0 processBlock - 250000 vs - 7PWei normalcase 2.4773 ms/op 2.7845 ms/op 0.89
phase0 processBlock - 250000 vs - 7PWei worstcase 29.370 ms/op 30.920 ms/op 0.95
altair processEth1Data - 250000 vs - 7PWei normalcase 472.26 us/op 499.87 us/op 0.94
getExpectedWithdrawals 250000 eb:1,eth1:1,we:0,wn:0,smpl:15 15.153 us/op 17.225 us/op 0.88
getExpectedWithdrawals 250000 eb:0.95,eth1:0.1,we:0.05,wn:0,smpl:219 37.433 us/op 96.998 us/op 0.39
getExpectedWithdrawals 250000 eb:0.95,eth1:0.3,we:0.05,wn:0,smpl:42 27.274 us/op 15.655 us/op 1.74
getExpectedWithdrawals 250000 eb:0.95,eth1:0.7,we:0.05,wn:0,smpl:18 12.622 us/op 9.9430 us/op 1.27
getExpectedWithdrawals 250000 eb:0.1,eth1:0.1,we:0,wn:0,smpl:1020 214.98 us/op 287.11 us/op 0.75
getExpectedWithdrawals 250000 eb:0.03,eth1:0.03,we:0,wn:0,smpl:11777 1.6979 ms/op 2.3363 ms/op 0.73
getExpectedWithdrawals 250000 eb:0.01,eth1:0.01,we:0,wn:0,smpl:16384 2.8727 ms/op 2.8221 ms/op 1.02
getExpectedWithdrawals 250000 eb:0,eth1:0,we:0,wn:0,smpl:16384 2.9177 ms/op 2.8476 ms/op 1.02
getExpectedWithdrawals 250000 eb:0,eth1:0,we:0,wn:0,nocache,smpl:16384 3.5486 ms/op 3.2481 ms/op 1.09
getExpectedWithdrawals 250000 eb:0,eth1:1,we:0,wn:0,smpl:16384 2.3184 ms/op 2.3481 ms/op 0.99
getExpectedWithdrawals 250000 eb:0,eth1:1,we:0,wn:0,nocache,smpl:16384 5.3075 ms/op 5.5503 ms/op 0.96
Tree 40 250000 create 326.67 ms/op 381.18 ms/op 0.86
Tree 40 250000 get(125000) 185.36 ns/op 214.27 ns/op 0.87
Tree 40 250000 set(125000) 964.27 ns/op 1.0639 us/op 0.91
Tree 40 250000 toArray() 18.717 ms/op 19.710 ms/op 0.95
Tree 40 250000 iterate all - toArray() + loop 18.877 ms/op 20.286 ms/op 0.93
Tree 40 250000 iterate all - get(i) 65.628 ms/op 71.670 ms/op 0.92
MutableVector 250000 create 19.961 ms/op 18.243 ms/op 1.09
MutableVector 250000 get(125000) 6.3750 ns/op 6.7540 ns/op 0.94
MutableVector 250000 set(125000) 251.32 ns/op 264.76 ns/op 0.95
MutableVector 250000 toArray() 3.6510 ms/op 3.4884 ms/op 1.05
MutableVector 250000 iterate all - toArray() + loop 3.9603 ms/op 3.9021 ms/op 1.01
MutableVector 250000 iterate all - get(i) 1.5047 ms/op 1.6354 ms/op 0.92
Array 250000 create 3.1069 ms/op 3.3386 ms/op 0.93
Array 250000 clone - spread 1.3138 ms/op 1.3772 ms/op 0.95
Array 250000 get(125000) 1.0530 ns/op 1.1390 ns/op 0.92
Array 250000 set(125000) 4.1220 ns/op 4.2810 ns/op 0.96
Array 250000 iterate all - loop 164.70 us/op 186.57 us/op 0.88
effectiveBalanceIncrements clone Uint8Array 300000 27.235 us/op 43.168 us/op 0.63
effectiveBalanceIncrements clone MutableVector 300000 387.00 ns/op 481.00 ns/op 0.80
effectiveBalanceIncrements rw all Uint8Array 300000 198.91 us/op 241.03 us/op 0.83
effectiveBalanceIncrements rw all MutableVector 300000 84.313 ms/op 109.23 ms/op 0.77
phase0 afterProcessEpoch - 250000 vs - 7PWei 110.69 ms/op 131.85 ms/op 0.84
phase0 beforeProcessEpoch - 250000 vs - 7PWei 56.817 ms/op 62.078 ms/op 0.92
altair processEpoch - mainnet_e81889 545.88 ms/op 599.07 ms/op 0.91
mainnet_e81889 - altair beforeProcessEpoch 84.316 ms/op 104.56 ms/op 0.81
mainnet_e81889 - altair processJustificationAndFinalization 15.720 us/op 33.639 us/op 0.47
mainnet_e81889 - altair processInactivityUpdates 5.2414 ms/op 9.3013 ms/op 0.56
mainnet_e81889 - altair processRewardsAndPenalties 89.693 ms/op 75.536 ms/op 1.19
mainnet_e81889 - altair processRegistryUpdates 3.4060 us/op 5.3300 us/op 0.64
mainnet_e81889 - altair processSlashings 891.00 ns/op 1.0830 us/op 0.82
mainnet_e81889 - altair processEth1DataReset 745.00 ns/op 947.00 ns/op 0.79
mainnet_e81889 - altair processEffectiveBalanceUpdates 2.5786 ms/op 2.0591 ms/op 1.25
mainnet_e81889 - altair processSlashingsReset 5.5260 us/op 5.7430 us/op 0.96
mainnet_e81889 - altair processRandaoMixesReset 5.8860 us/op 5.1950 us/op 1.13
mainnet_e81889 - altair processHistoricalRootsUpdate 1.0280 us/op 1.3070 us/op 0.79
mainnet_e81889 - altair processParticipationFlagUpdates 3.0460 us/op 4.1830 us/op 0.73
mainnet_e81889 - altair processSyncCommitteeUpdates 770.00 ns/op 1.2070 us/op 0.64
mainnet_e81889 - altair afterProcessEpoch 115.67 ms/op 160.60 ms/op 0.72
capella processEpoch - mainnet_e217614 2.4719 s/op 3.8050 s/op 0.65
mainnet_e217614 - capella beforeProcessEpoch 492.86 ms/op 910.47 ms/op 0.54
mainnet_e217614 - capella processJustificationAndFinalization 20.014 us/op 40.597 us/op 0.49
mainnet_e217614 - capella processInactivityUpdates 17.316 ms/op 40.853 ms/op 0.42
mainnet_e217614 - capella processRewardsAndPenalties 566.71 ms/op 1.0767 s/op 0.53
mainnet_e217614 - capella processRegistryUpdates 26.140 us/op 40.287 us/op 0.65
mainnet_e217614 - capella processSlashings 602.00 ns/op 1.2280 us/op 0.49
mainnet_e217614 - capella processEth1DataReset 562.00 ns/op 979.00 ns/op 0.57
mainnet_e217614 - capella processEffectiveBalanceUpdates 13.921 ms/op 7.4595 ms/op 1.87
mainnet_e217614 - capella processSlashingsReset 3.2210 us/op 8.3230 us/op 0.39
mainnet_e217614 - capella processRandaoMixesReset 4.2720 us/op 13.754 us/op 0.31
mainnet_e217614 - capella processHistoricalRootsUpdate 881.00 ns/op 1.6280 us/op 0.54
mainnet_e217614 - capella processParticipationFlagUpdates 2.1910 us/op 6.1610 us/op 0.36
mainnet_e217614 - capella afterProcessEpoch 316.89 ms/op 643.05 ms/op 0.49
phase0 processEpoch - mainnet_e58758 622.28 ms/op 601.25 ms/op 1.03
mainnet_e58758 - phase0 beforeProcessEpoch 188.43 ms/op 179.14 ms/op 1.05
mainnet_e58758 - phase0 processJustificationAndFinalization 24.002 us/op 23.191 us/op 1.03
mainnet_e58758 - phase0 processRewardsAndPenalties 70.070 ms/op 89.238 ms/op 0.79
mainnet_e58758 - phase0 processRegistryUpdates 12.899 us/op 22.553 us/op 0.57
mainnet_e58758 - phase0 processSlashings 548.00 ns/op 1.2240 us/op 0.45
mainnet_e58758 - phase0 processEth1DataReset 575.00 ns/op 1.2140 us/op 0.47
mainnet_e58758 - phase0 processEffectiveBalanceUpdates 1.3991 ms/op 2.0296 ms/op 0.69
mainnet_e58758 - phase0 processSlashingsReset 3.1790 us/op 6.3360 us/op 0.50
mainnet_e58758 - phase0 processRandaoMixesReset 7.1340 us/op 15.414 us/op 0.46
mainnet_e58758 - phase0 processHistoricalRootsUpdate 915.00 ns/op 2.0920 us/op 0.44
mainnet_e58758 - phase0 processParticipationRecordUpdates 8.1520 us/op 7.6050 us/op 1.07
mainnet_e58758 - phase0 afterProcessEpoch 111.16 ms/op 171.18 ms/op 0.65
phase0 processEffectiveBalanceUpdates - 250000 normalcase 1.6547 ms/op 2.2512 ms/op 0.74
phase0 processEffectiveBalanceUpdates - 250000 worstcase 0.5 2.2474 ms/op 3.7693 ms/op 0.60
altair processInactivityUpdates - 250000 normalcase 34.047 ms/op 47.850 ms/op 0.71
altair processInactivityUpdates - 250000 worstcase 40.459 ms/op 57.509 ms/op 0.70
phase0 processRegistryUpdates - 250000 normalcase 11.825 us/op 19.589 us/op 0.60
phase0 processRegistryUpdates - 250000 badcase_full_deposits 564.25 us/op 815.68 us/op 0.69
phase0 processRegistryUpdates - 250000 worstcase 0.5 203.48 ms/op 350.19 ms/op 0.58
altair processRewardsAndPenalties - 250000 normalcase 84.067 ms/op 128.19 ms/op 0.66
altair processRewardsAndPenalties - 250000 worstcase 68.241 ms/op 141.87 ms/op 0.48
phase0 getAttestationDeltas - 250000 normalcase 13.924 ms/op 19.168 ms/op 0.73
phase0 getAttestationDeltas - 250000 worstcase 15.404 ms/op 18.380 ms/op 0.84
phase0 processSlashings - 250000 worstcase 130.14 us/op 151.19 us/op 0.86
altair processSyncCommitteeUpdates - 250000 206.77 ms/op 360.98 ms/op 0.57
BeaconState.hashTreeRoot - No change 721.00 ns/op 1.0510 us/op 0.69
BeaconState.hashTreeRoot - 1 full validator 150.62 us/op 272.93 us/op 0.55
BeaconState.hashTreeRoot - 32 full validator 1.6295 ms/op 2.6845 ms/op 0.61
BeaconState.hashTreeRoot - 512 full validator 19.997 ms/op 26.730 ms/op 0.75
BeaconState.hashTreeRoot - 1 validator.effectiveBalance 219.30 us/op 277.42 us/op 0.79
BeaconState.hashTreeRoot - 32 validator.effectiveBalance 2.7143 ms/op 3.9716 ms/op 0.68
BeaconState.hashTreeRoot - 512 validator.effectiveBalance 36.255 ms/op 55.088 ms/op 0.66
BeaconState.hashTreeRoot - 1 balances 157.64 us/op 272.57 us/op 0.58
BeaconState.hashTreeRoot - 32 balances 1.4249 ms/op 2.0705 ms/op 0.69
BeaconState.hashTreeRoot - 512 balances 14.523 ms/op 15.385 ms/op 0.94
BeaconState.hashTreeRoot - 250000 balances 255.44 ms/op 300.66 ms/op 0.85
aggregationBits - 2048 els - zipIndexesInBitList 36.605 us/op 51.637 us/op 0.71
byteArrayEquals 32 78.191 ns/op 129.92 ns/op 0.60
Buffer.compare 32 58.470 ns/op 97.843 ns/op 0.60
byteArrayEquals 1024 2.1147 us/op 2.5461 us/op 0.83
Buffer.compare 1024 72.146 ns/op 107.50 ns/op 0.67
byteArrayEquals 16384 33.785 us/op 43.840 us/op 0.77
Buffer.compare 16384 270.01 ns/op 434.59 ns/op 0.62
byteArrayEquals 123687377 243.37 ms/op 333.89 ms/op 0.73
Buffer.compare 123687377 6.0411 ms/op 9.0505 ms/op 0.67
byteArrayEquals 32 - diff last byte 71.744 ns/op 137.83 ns/op 0.52
Buffer.compare 32 - diff last byte 54.878 ns/op 104.17 ns/op 0.53
byteArrayEquals 1024 - diff last byte 1.9743 us/op 3.2307 us/op 0.61
Buffer.compare 1024 - diff last byte 69.756 ns/op 99.016 ns/op 0.70
byteArrayEquals 16384 - diff last byte 31.510 us/op 40.872 us/op 0.77
Buffer.compare 16384 - diff last byte 261.23 ns/op 290.68 ns/op 0.90
byteArrayEquals 123687377 - diff last byte 237.24 ms/op 329.83 ms/op 0.72
Buffer.compare 123687377 - diff last byte 6.2403 ms/op 8.8040 ms/op 0.71
byteArrayEquals 32 - random bytes 5.0770 ns/op 7.0580 ns/op 0.72
Buffer.compare 32 - random bytes 58.729 ns/op 80.322 ns/op 0.73
byteArrayEquals 1024 - random bytes 4.9870 ns/op 8.0610 ns/op 0.62
Buffer.compare 1024 - random bytes 58.052 ns/op 99.319 ns/op 0.58
byteArrayEquals 16384 - random bytes 4.9850 ns/op 8.6240 ns/op 0.58
Buffer.compare 16384 - random bytes 58.335 ns/op 101.94 ns/op 0.57
byteArrayEquals 123687377 - random bytes 8.0100 ns/op 12.120 ns/op 0.66
Buffer.compare 123687377 - random bytes 61.110 ns/op 103.32 ns/op 0.59
regular array get 100000 times 42.516 us/op 72.116 us/op 0.59
wrappedArray get 100000 times 42.612 us/op 57.728 us/op 0.74
arrayWithProxy get 100000 times 14.384 ms/op 18.671 ms/op 0.77
ssz.Root.equals 51.364 ns/op 81.232 ns/op 0.63
byteArrayEquals 50.685 ns/op 89.944 ns/op 0.56
Buffer.compare 10.371 ns/op 18.338 ns/op 0.57
shuffle list - 16384 els 8.4060 ms/op 10.981 ms/op 0.77
shuffle list - 250000 els 122.99 ms/op 178.82 ms/op 0.69
processSlot - 1 slots 20.290 us/op 24.851 us/op 0.82
processSlot - 32 slots 3.9100 ms/op 5.3111 ms/op 0.74
getEffectiveBalanceIncrementsZeroInactive - 250000 vs - 7PWei 62.727 ms/op 73.044 ms/op 0.86
getCommitteeAssignments - req 1 vs - 250000 vc 2.6928 ms/op 4.2270 ms/op 0.64
getCommitteeAssignments - req 100 vs - 250000 vc 3.8964 ms/op 5.2073 ms/op 0.75
getCommitteeAssignments - req 1000 vs - 250000 vc 4.1984 ms/op 6.3540 ms/op 0.66
findModifiedValidators - 10000 modified validators 360.05 ms/op 499.09 ms/op 0.72
findModifiedValidators - 1000 modified validators 208.24 ms/op 269.74 ms/op 0.77
findModifiedValidators - 100 modified validators 172.66 ms/op 231.46 ms/op 0.75
findModifiedValidators - 10 modified validators 194.48 ms/op 252.66 ms/op 0.77
findModifiedValidators - 1 modified validators 186.35 ms/op 278.10 ms/op 0.67
findModifiedValidators - no difference 195.53 ms/op 271.88 ms/op 0.72
compare ViewDUs 4.8359 s/op 5.7213 s/op 0.85
compare each validator Uint8Array 1.9708 s/op 2.1696 s/op 0.91
compare ViewDU to Uint8Array 1.1628 s/op 1.5887 s/op 0.73
migrate state 1000000 validators, 24 modified, 0 new 825.79 ms/op 984.51 ms/op 0.84
migrate state 1000000 validators, 1700 modified, 1000 new 1.2824 s/op 1.4267 s/op 0.90
migrate state 1000000 validators, 3400 modified, 2000 new 1.6213 s/op 1.4867 s/op 1.09
migrate state 1500000 validators, 24 modified, 0 new 891.52 ms/op 924.52 ms/op 0.96
migrate state 1500000 validators, 1700 modified, 1000 new 1.2681 s/op 1.2331 s/op 1.03
migrate state 1500000 validators, 3400 modified, 2000 new 1.6308 s/op 1.6931 s/op 0.96
RootCache.getBlockRootAtSlot - 250000 vs - 7PWei 4.8700 ns/op 5.8800 ns/op 0.83
state getBlockRootAtSlot - 250000 vs - 7PWei 653.02 ns/op 980.59 ns/op 0.67
computeProposers - vc 250000 9.7322 ms/op 14.842 ms/op 0.66
computeEpochShuffling - vc 250000 130.06 ms/op 248.15 ms/op 0.52
getNextSyncCommittee - vc 250000 182.87 ms/op 275.81 ms/op 0.66
computeSigningRoot for AttestationData 35.439 us/op 50.665 us/op 0.70
hash AttestationData serialized data then Buffer.toString(base64) 2.3770 us/op 3.6679 us/op 0.65
toHexString serialized data 1.1327 us/op 2.2941 us/op 0.49
Buffer.toString(base64) 244.29 ns/op 414.19 ns/op 0.59

by benchmarkbot/action

Comment thread packages/beacon-node/src/chain/bls/multithread/jobItem.ts Outdated
Comment thread packages/beacon-node/src/chain/bls/multithread/jobItem.ts
Comment thread packages/beacon-node/src/chain/bls/maybeBatch.ts
Comment thread packages/beacon-node/src/chain/bls/multithread/index.ts Outdated
Comment thread packages/beacon-node/src/chain/bls/multithread/index.ts Outdated
@matthewkeil
matthewkeil force-pushed the mkeil/napi-rebuild-2-rebase-unstable-final-bls branch from 7fb847a to 8dda2ff Compare April 12, 2024 11:51
} from "./jobItem.js";
import {asyncVerifyManySignatureSets} from "./verifyManySignatureSets.js";

// defaultPoolSize should return core count - 1 to keep main thread on core. We

@nflaig nflaig Apr 13, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am wondering if we should always go with logical CPU count - 1 or have a max cap at some point, e.g. does it make sense to have pool size of 64 or even 128?

Same goes for other worker pools we use

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have it sized at number of cores. There is a fair amount of idle time on the main and network thread so the kernel is switching those out for bls in down times.

I need to update that note.

I tried with cores - 2 at first to keep main and network thread on at all times but there was performance degradation. Tried cores -1 also but found the better but similar. Best results were sizing the pool at the number of cores and letting the kernel do the management. Can fine tune further though I think

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

But regarding my question, is there an upper bound to the pool where it just does not make sense to spawn more workers?

I tried with cores - 2 at first to keep main and network thread on at all times but there was performance degradation.

I am surprised this makes a difference, because we kinda assume here that Lodestar is the only process running on the server which for most setups is not the case.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would assume more than the number of cores is generally bad. So as an upper limit I tried to stick with number of cores.

The bls work is very math heavy and the assembly portions of the code (synchronous functions) stay on thread for most of the execution time. Attempting to do more parallel calls than there are cores will net a reduction of overall performance because of the time swapping thread context (moving active thread to wait state while another thread is being executed by a core). Similar to how we are moving the main thread and worker thread to wait while bls runs. The kernel knows when a process is idle though and optimizes for this so that processes that are active get run and ones that are idle can wait for other threads to run.

There are some other nuances to that though. Multi-threading allows a core to have more than one "active" thread that can execute at any give time although this is not a standardized thing. And how/when the thread is "active" also is non-standard. Each manf. has a different core/thread ratio and different way of handling what is actually running assembly at a given time. I would guess that we can tune the number of threads up slightly but how much will be empirical, and results will be specific to the machine they were tested with.

As for the part that the EL is also running on the box. This is a good point. My evidence for setting the threadpool size was empirical. Ran the same branch with a few different values on separate feature groups. And was rough at best but the results were pretty obvious at the time however they did not get persisted anywhere other that me going "cores = threads is better" and I went with it. There was a lot of variables floating around at the time so i just simply went with what looked best. At this point, now that we have narrowed a lot down and have a much more stable set of factors it will be a good idea to do a more methodical approach to setting this value.

Also will be a good idea to include EL performance in the calculation for setting. Not sure how to rate that side's efficiency but should be easy enough with a bit of digging. Will def add this to the to-do list

@matthewkeil matthewkeil Apr 16, 2024

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@nflaig as an example I found on OSX that cores = threads (ie 10 on my M1) but on our hzax41 boxes the threads = 2 x cores. Generally when requesting the "number of cores" or "how much multithreading" is available the number of logical cores is what is reported.

Linux - Cloud VM

> user@mainnet-hzax41:~$ awk -F': ' '/cpu cores/{print $2;exit}' /proc/cpuinfo
6
> user@mainnet-hzax41:~$ grep -c ^processor /proc/cpuinfo
12
console.log(require(`os`).availableParallelism())
// > 12

OSX - M1 (some irrelevant lines removed)

❯ sysctl -a | grep cpu
hw.ncpu: 10
hw.activecpu: 10
hw.physicalcpu: 10
hw.physicalcpu_max: 10
hw.logicalcpu: 10
hw.logicalcpu_max: 10
machdep.cpu.cores_per_package: 10
machdep.cpu.core_count: 10
machdep.cpu.logical_per_package: 10
machdep.cpu.thread_count: 10
machdep.cpu.brand_string: Apple M1 Max
console.log(require(`os`).availableParallelism())
// > 10

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@nflaig as an example I found on OSX that cores = threads (ie 10 on my M1) but on our hzax41 boxes the threads = 2 x cores. Generally when requesting the "number of cores" or "how much multithreading" is available the number of logical cores is what is reported.

Also be aware of docker, often times I saw logical CPUs to equal physical cores there.

@matthewkeil matthewkeil Apr 16, 2024

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I just noticed that at least for worker_threads, it's a noticeable overhead at startup, see issue during tests #5608 but I am not sue this matters. That's why I was asking because you have a better idea about this then I do

worker_thread != kernel thread though.... for a worker thread a whole JS environment is spun up and that does take a considerable amount of time. kernel threads are not the same... the time is not 0 but its very very very low (likely milliseconds). Not sure how long though but I can tell you that starting this branch with libuv it goes to the next step almost immediately (not sure exactly where in the flow the workers are started though in practice. it could be even before the bls class gets instantiated when fs happens or it could be when the first bls work gets queued during sync but not sure honestly) and starting with workers there is maybe a 10-15 second delay to start all the nodejs worker_threads

Found this on SO:
https://stackoverflow.com/questions/18274217/how-long-does-thread-creation-and-termination-take-under-windows

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@nflaig as an example I found on OSX that cores = threads (ie 10 on my M1) but on our hzax41 boxes the threads = 2 x cores. Generally when requesting the "number of cores" or "how much multithreading" is available the number of logical cores is what is reported.

Also be aware of docker, often times I saw logical CPUs to equal physical cores there.

I would guess, but need to confirm 100%, that docker would read the logical cpu count and report that to the container as the physical and logical count. The docker runtime has a setting to throttle how many "CPU's" are available to the daemon. I would assume they are using the soft/logical limit for that similar to how all software consumes the underlying system resources.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would guess, but need to confirm 100%, that docker would read the logical cpu count and report that to the container as the physical and logical count.

That was something I noticed when working on system metrics

const cpu = await system.cpu();
// Note: inside container this might be inaccurate as
// physicalCores in some cases is the count of logical CPU cores
this.cpuCores = cpu.physicalCores;
this.cpuThreads = cpu.cores;

Looking at the metrics from my server, it seems to be still the case

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

worker_thread != kernel thread though

yeah, that's what I thought as well, it might not matter for libuv threads but we should probably consider an upper bound for any other worker pools, although might not be relevant if we anyhow replace the BLS worker implementation, and it's not that problematic for keystore decryption as those are limited by number of keys already and will be removed once decryption is completed.

Comment thread packages/beacon-node/src/chain/bls/multithread/index.ts Outdated
@twoeths

twoeths commented Apr 16, 2024

Copy link
Copy Markdown
Member

latest profile show a lot of time used for timer

Screenshot 2024-04-16 at 08 51 56

I think that's also the reason why gc is really high, up to 60% on 1k node

feat2_1k_main_thread_2024-04-16T01:34:42.576Z.cpuprofile.zip

we really need to stop runJob() when there is nothing in the jobs queue

@matthewkeil

Copy link
Copy Markdown
Member Author

closed in favor of #6616

@matthewkeil
matthewkeil deleted the mkeil/napi-rebuild-2-rebase-unstable-final-bls branch May 23, 2024 10:03
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants