Skip to content

Skip deserializing gossip attestation messages by caching AttestationData - #5363

Merged
wemeetagain merged 8 commits into
unstablefrom
tuyen/cache_attestation_data
Apr 17, 2023
Merged

Skip deserializing gossip attestation messages by caching AttestationData#5363
wemeetagain merged 8 commits into
unstablefrom
tuyen/cache_attestation_data

Conversation

@twoeths

@twoeths twoeths commented Apr 14, 2023

Copy link
Copy Markdown
Member

Motivation

  • Improve gossip validation flow for beacon_attestation topic
  • We deserialize Attestation for every attestation gossip message but a lot of them have same AttestationData while:
    • aggregationBits can be extracted from ssz bytes
    • signature can also be extracted from ssz bytes
  • Caching AttestationData does not take memory because according to metrics, we process up to 15k-20k attestations per slot while this cache only have 600 items max
  • Fix NetworkProcessor.onPendingGossipsubMessage() to not push to the queue if unknown block root attestations

Description

  • Skip deserializing gossip attestation messages most of the time
  • Add AttestationData to seen cache, used for beacon_attestation topic only
  • gossipHandlerFn now accept ssz bytes instead of ssz object. For beacon_attestation we don't do the deserialization and delegate to the validateAttestation function to do that. For all other topics we deserialize to ssz objects and pass to validation functions (unchange)
  • Modify validateAttestation:
    • For api attestation, still accept Attestation object
    • For gossip attestation, accept ssz bytes. Only do the deserialization the 1st time AttestationData is not cached. After that reuse the cached AttestationData along with extracted aggregationBits and signature from ssz bytes.

part of #5352 #5353

Testing in progress on feat1, feat2

@github-actions

github-actions Bot commented Apr 14, 2023

Copy link
Copy Markdown
Contributor

Performance Report

✔️ no performance regression detected

Full benchmark results
Benchmark suite Current: 692bcff Previous: 42ecf97 Ratio
getPubkeys - index2pubkey - req 1000 vs - 250000 vc 872.67 us/op 879.43 us/op 0.99
getPubkeys - validatorsArr - req 1000 vs - 250000 vc 52.164 us/op 44.016 us/op 1.19
BLS verify - blst-native 1.2633 ms/op 1.1868 ms/op 1.06
BLS verifyMultipleSignatures 3 - blst-native 2.5893 ms/op 2.4144 ms/op 1.07
BLS verifyMultipleSignatures 8 - blst-native 5.6489 ms/op 5.1871 ms/op 1.09
BLS verifyMultipleSignatures 32 - blst-native 19.975 ms/op 18.766 ms/op 1.06
BLS aggregatePubkeys 32 - blst-native 27.126 us/op 25.342 us/op 1.07
BLS aggregatePubkeys 128 - blst-native 105.53 us/op 97.819 us/op 1.08
getAttestationsForBlock 58.538 ms/op 51.986 ms/op 1.13
isKnown best case - 1 super set check 271.00 ns/op 255.00 ns/op 1.06
isKnown normal case - 2 super set checks 262.00 ns/op 249.00 ns/op 1.05
isKnown worse case - 16 super set checks 264.00 ns/op 250.00 ns/op 1.06
CheckpointStateCache - add get delete 5.8290 us/op 4.7990 us/op 1.21
validate gossip signedAggregateAndProof - struct 2.8845 ms/op 2.6813 ms/op 1.08
validate gossip attestation - struct 1.3755 ms/op 1.2911 ms/op 1.07
pickEth1Vote - no votes 1.4238 ms/op 1.2075 ms/op 1.18
pickEth1Vote - max votes 11.521 ms/op 10.178 ms/op 1.13
pickEth1Vote - Eth1Data hashTreeRoot value x2048 9.6217 ms/op 8.7238 ms/op 1.10
pickEth1Vote - Eth1Data hashTreeRoot tree x2048 16.140 ms/op 14.579 ms/op 1.11
pickEth1Vote - Eth1Data fastSerialize value x2048 792.05 us/op 617.60 us/op 1.28
pickEth1Vote - Eth1Data fastSerialize tree x2048 8.8887 ms/op 7.7105 ms/op 1.15
bytes32 toHexString 648.00 ns/op 473.00 ns/op 1.37
bytes32 Buffer.toString(hex) 411.00 ns/op 342.00 ns/op 1.20
bytes32 Buffer.toString(hex) from Uint8Array 638.00 ns/op 553.00 ns/op 1.15
bytes32 Buffer.toString(hex) + 0x 423.00 ns/op 345.00 ns/op 1.23
Object access 1 prop 0.21000 ns/op 0.15700 ns/op 1.34
Map access 1 prop 0.16800 ns/op 0.15200 ns/op 1.11
Object get x1000 7.0290 ns/op 6.2530 ns/op 1.12
Map get x1000 0.60000 ns/op 0.58700 ns/op 1.02
Object set x1000 72.061 ns/op 49.909 ns/op 1.44
Map set x1000 53.925 ns/op 41.890 ns/op 1.29
Return object 10000 times 0.26230 ns/op 0.22720 ns/op 1.15
Throw Error 10000 times 4.4017 us/op 3.9663 us/op 1.11
fastMsgIdFn sha256 / 200 bytes 3.6840 us/op 3.3320 us/op 1.11
fastMsgIdFn h32 xxhash / 200 bytes 316.00 ns/op 272.00 ns/op 1.16
fastMsgIdFn h64 xxhash / 200 bytes 459.00 ns/op 375.00 ns/op 1.22
fastMsgIdFn sha256 / 1000 bytes 12.212 us/op 11.325 us/op 1.08
fastMsgIdFn h32 xxhash / 1000 bytes 453.00 ns/op 403.00 ns/op 1.12
fastMsgIdFn h64 xxhash / 1000 bytes 553.00 ns/op 444.00 ns/op 1.25
fastMsgIdFn sha256 / 10000 bytes 106.56 us/op 100.73 us/op 1.06
fastMsgIdFn h32 xxhash / 10000 bytes 2.0510 us/op 1.8450 us/op 1.11
fastMsgIdFn h64 xxhash / 10000 bytes 1.4930 us/op 1.2790 us/op 1.17
enrSubnets - fastDeserialize 64 bits 1.8000 us/op 1.2360 us/op 1.46
enrSubnets - ssz BitVector 64 bits 620.00 ns/op 459.00 ns/op 1.35
enrSubnets - fastDeserialize 4 bits 210.00 ns/op 160.00 ns/op 1.31
enrSubnets - ssz BitVector 4 bits 617.00 ns/op 478.00 ns/op 1.29
prioritizePeers score -10:0 att 32-0.1 sync 2-0 116.68 us/op 104.19 us/op 1.12
prioritizePeers score 0:0 att 32-0.25 sync 2-0.25 157.79 us/op 127.27 us/op 1.24
prioritizePeers score 0:0 att 32-0.5 sync 2-0.5 198.51 us/op 163.28 us/op 1.22
prioritizePeers score 0:0 att 64-0.75 sync 4-0.75 366.28 us/op 301.10 us/op 1.22
prioritizePeers score 0:0 att 64-1 sync 4-1 442.73 us/op 347.10 us/op 1.28
array of 16000 items push then shift 1.7246 us/op 1.5470 us/op 1.11
LinkedList of 16000 items push then shift 9.3790 ns/op 8.3660 ns/op 1.12
array of 16000 items push then pop 116.27 ns/op 73.759 ns/op 1.58
LinkedList of 16000 items push then pop 9.5920 ns/op 8.0890 ns/op 1.19
array of 24000 items push then shift 2.4912 us/op 2.2436 us/op 1.11
LinkedList of 24000 items push then shift 10.216 ns/op 8.3700 ns/op 1.22
array of 24000 items push then pop 89.194 ns/op 72.860 ns/op 1.22
LinkedList of 24000 items push then pop 9.5180 ns/op 8.1220 ns/op 1.17
intersect bitArray bitLen 8 13.696 ns/op 12.698 ns/op 1.08
intersect array and set length 8 91.671 ns/op 73.733 ns/op 1.24
intersect bitArray bitLen 128 45.541 ns/op 42.095 ns/op 1.08
intersect array and set length 128 1.2986 us/op 1.0029 us/op 1.29
Buffer.concat 32 items 3.0970 us/op 2.8690 us/op 1.08
Uint8Array.set 32 items 2.7070 us/op 2.8310 us/op 0.96
pass gossip attestations to forkchoice per slot 3.2053 ms/op 2.6937 ms/op 1.19
computeDeltas 3.4036 ms/op 3.2391 ms/op 1.05
computeProposerBoostScoreFromBalances 1.8696 ms/op 1.7650 ms/op 1.06
altair processAttestation - 250000 vs - 7PWei normalcase 3.3579 ms/op 2.2279 ms/op 1.51
altair processAttestation - 250000 vs - 7PWei worstcase 4.4003 ms/op 3.6420 ms/op 1.21
altair processAttestation - setStatus - 1/6 committees join 151.80 us/op 143.19 us/op 1.06
altair processAttestation - setStatus - 1/3 committees join 287.97 us/op 280.22 us/op 1.03
altair processAttestation - setStatus - 1/2 committees join 385.30 us/op 373.49 us/op 1.03
altair processAttestation - setStatus - 2/3 committees join 478.66 us/op 461.95 us/op 1.04
altair processAttestation - setStatus - 4/5 committees join 678.52 us/op 653.16 us/op 1.04
altair processAttestation - setStatus - 100% committees join 790.58 us/op 757.24 us/op 1.04
altair processBlock - 250000 vs - 7PWei normalcase 19.041 ms/op 18.447 ms/op 1.03
altair processBlock - 250000 vs - 7PWei normalcase hashState 24.984 ms/op 25.932 ms/op 0.96
altair processBlock - 250000 vs - 7PWei worstcase 47.985 ms/op 52.179 ms/op 0.92
altair processBlock - 250000 vs - 7PWei worstcase hashState 67.697 ms/op 67.562 ms/op 1.00
phase0 processBlock - 250000 vs - 7PWei normalcase 2.5494 ms/op 2.2034 ms/op 1.16
phase0 processBlock - 250000 vs - 7PWei worstcase 31.137 ms/op 29.882 ms/op 1.04
altair processEth1Data - 250000 vs - 7PWei normalcase 579.27 us/op 529.52 us/op 1.09
vc - 250000 eb 1 eth1 1 we 0 wn 0 - smpl 15 9.3870 us/op 8.2940 us/op 1.13
vc - 250000 eb 0.95 eth1 0.1 we 0.05 wn 0 - smpl 219 28.573 us/op 26.927 us/op 1.06
vc - 250000 eb 0.95 eth1 0.3 we 0.05 wn 0 - smpl 42 11.156 us/op 11.188 us/op 1.00
vc - 250000 eb 0.95 eth1 0.7 we 0.05 wn 0 - smpl 18 8.9300 us/op 8.4570 us/op 1.06
vc - 250000 eb 0.1 eth1 0.1 we 0 wn 0 - smpl 1020 116.11 us/op 99.509 us/op 1.17
vc - 250000 eb 0.03 eth1 0.03 we 0 wn 0 - smpl 11777 690.60 us/op 658.01 us/op 1.05
vc - 250000 eb 0.01 eth1 0.01 we 0 wn 0 - smpl 16384 928.38 us/op 918.66 us/op 1.01
vc - 250000 eb 0 eth1 0 we 0 wn 0 - smpl 16384 1.0197 ms/op 906.88 us/op 1.12
vc - 250000 eb 0 eth1 0 we 0 wn 0 nocache - smpl 16384 2.9928 ms/op 2.5046 ms/op 1.19
vc - 250000 eb 0 eth1 1 we 0 wn 0 - smpl 16384 1.7506 ms/op 1.5665 ms/op 1.12
vc - 250000 eb 0 eth1 1 we 0 wn 0 nocache - smpl 16384 4.5350 ms/op 4.0918 ms/op 1.11
Tree 40 250000 create 345.18 ms/op 387.86 ms/op 0.89
Tree 40 250000 get(125000) 196.54 ns/op 201.59 ns/op 0.97
Tree 40 250000 set(125000) 1.0711 us/op 939.25 ns/op 1.14
Tree 40 250000 toArray() 24.544 ms/op 22.956 ms/op 1.07
Tree 40 250000 iterate all - toArray() + loop 24.170 ms/op 21.874 ms/op 1.10
Tree 40 250000 iterate all - get(i) 80.893 ms/op 74.149 ms/op 1.09
MutableVector 250000 create 11.545 ms/op 10.606 ms/op 1.09
MutableVector 250000 get(125000) 6.5750 ns/op 6.9210 ns/op 0.95
MutableVector 250000 set(125000) 284.15 ns/op 255.04 ns/op 1.11
MutableVector 250000 toArray() 4.1636 ms/op 3.5254 ms/op 1.18
MutableVector 250000 iterate all - toArray() + loop 4.5588 ms/op 3.5547 ms/op 1.28
MutableVector 250000 iterate all - get(i) 1.5935 ms/op 1.5610 ms/op 1.02
Array 250000 create 4.1125 ms/op 2.8978 ms/op 1.42
Array 250000 clone - spread 1.3198 ms/op 1.1546 ms/op 1.14
Array 250000 get(125000) 0.66300 ns/op 0.57100 ns/op 1.16
Array 250000 set(125000) 0.72100 ns/op 0.64700 ns/op 1.11
Array 250000 iterate all - loop 96.052 us/op 94.497 us/op 1.02
effectiveBalanceIncrements clone Uint8Array 300000 47.992 us/op 30.317 us/op 1.58
effectiveBalanceIncrements clone MutableVector 300000 411.00 ns/op 338.00 ns/op 1.22
effectiveBalanceIncrements rw all Uint8Array 300000 186.60 us/op 168.38 us/op 1.11
effectiveBalanceIncrements rw all MutableVector 300000 122.28 ms/op 79.430 ms/op 1.54
phase0 afterProcessEpoch - 250000 vs - 7PWei 122.87 ms/op 115.24 ms/op 1.07
phase0 beforeProcessEpoch - 250000 vs - 7PWei 44.139 ms/op 40.067 ms/op 1.10
altair processEpoch - mainnet_e81889 364.75 ms/op 295.79 ms/op 1.23
mainnet_e81889 - altair beforeProcessEpoch 73.584 ms/op 50.289 ms/op 1.46
mainnet_e81889 - altair processJustificationAndFinalization 20.901 us/op 17.940 us/op 1.17
mainnet_e81889 - altair processInactivityUpdates 6.5655 ms/op 5.5569 ms/op 1.18
mainnet_e81889 - altair processRewardsAndPenalties 52.370 ms/op 70.834 ms/op 0.74
mainnet_e81889 - altair processRegistryUpdates 3.0900 us/op 2.3340 us/op 1.32
mainnet_e81889 - altair processSlashings 726.00 ns/op 451.00 ns/op 1.61
mainnet_e81889 - altair processEth1DataReset 1.0570 us/op 464.00 ns/op 2.28
mainnet_e81889 - altair processEffectiveBalanceUpdates 1.4442 ms/op 1.2477 ms/op 1.16
mainnet_e81889 - altair processSlashingsReset 5.2820 us/op 4.3430 us/op 1.22
mainnet_e81889 - altair processRandaoMixesReset 8.8500 us/op 5.7690 us/op 1.53
mainnet_e81889 - altair processHistoricalRootsUpdate 1.2020 us/op 674.00 ns/op 1.78
mainnet_e81889 - altair processParticipationFlagUpdates 3.4120 us/op 2.7830 us/op 1.23
mainnet_e81889 - altair processSyncCommitteeUpdates 766.00 ns/op 610.00 ns/op 1.26
mainnet_e81889 - altair afterProcessEpoch 135.42 ms/op 126.83 ms/op 1.07
phase0 processEpoch - mainnet_e58758 379.58 ms/op 352.17 ms/op 1.08
mainnet_e58758 - phase0 beforeProcessEpoch 145.01 ms/op 136.99 ms/op 1.06
mainnet_e58758 - phase0 processJustificationAndFinalization 20.650 us/op 23.756 us/op 0.87
mainnet_e58758 - phase0 processRewardsAndPenalties 60.503 ms/op 63.144 ms/op 0.96
mainnet_e58758 - phase0 processRegistryUpdates 9.9260 us/op 8.7830 us/op 1.13
mainnet_e58758 - phase0 processSlashings 710.00 ns/op 519.00 ns/op 1.37
mainnet_e58758 - phase0 processEth1DataReset 1.0280 us/op 804.00 ns/op 1.28
mainnet_e58758 - phase0 processEffectiveBalanceUpdates 1.2085 ms/op 1.1123 ms/op 1.09
mainnet_e58758 - phase0 processSlashingsReset 6.8910 us/op 3.6690 us/op 1.88
mainnet_e58758 - phase0 processRandaoMixesReset 5.0380 us/op 4.7840 us/op 1.05
mainnet_e58758 - phase0 processHistoricalRootsUpdate 736.00 ns/op 611.00 ns/op 1.20
mainnet_e58758 - phase0 processParticipationRecordUpdates 9.9790 us/op 4.1110 us/op 2.43
mainnet_e58758 - phase0 afterProcessEpoch 98.228 ms/op 100.29 ms/op 0.98
phase0 processEffectiveBalanceUpdates - 250000 normalcase 1.2854 ms/op 1.2598 ms/op 1.02
phase0 processEffectiveBalanceUpdates - 250000 worstcase 0.5 1.5327 ms/op 1.5148 ms/op 1.01
altair processInactivityUpdates - 250000 normalcase 27.863 ms/op 26.644 ms/op 1.05
altair processInactivityUpdates - 250000 worstcase 29.520 ms/op 28.051 ms/op 1.05
phase0 processRegistryUpdates - 250000 normalcase 7.8410 us/op 6.6410 us/op 1.18
phase0 processRegistryUpdates - 250000 badcase_full_deposits 278.18 us/op 216.22 us/op 1.29
phase0 processRegistryUpdates - 250000 worstcase 0.5 127.65 ms/op 118.26 ms/op 1.08
altair processRewardsAndPenalties - 250000 normalcase 64.787 ms/op 66.574 ms/op 0.97
altair processRewardsAndPenalties - 250000 worstcase 72.485 ms/op 70.207 ms/op 1.03
phase0 getAttestationDeltas - 250000 normalcase 8.2806 ms/op 6.5894 ms/op 1.26
phase0 getAttestationDeltas - 250000 worstcase 8.9292 ms/op 6.5362 ms/op 1.37
phase0 processSlashings - 250000 worstcase 4.1555 ms/op 3.3369 ms/op 1.25
altair processSyncCommitteeUpdates - 250000 191.72 ms/op 171.01 ms/op 1.12
BeaconState.hashTreeRoot - No change 273.00 ns/op 259.00 ns/op 1.05
BeaconState.hashTreeRoot - 1 full validator 54.373 us/op 52.978 us/op 1.03
BeaconState.hashTreeRoot - 32 full validator 560.98 us/op 483.49 us/op 1.16
BeaconState.hashTreeRoot - 512 full validator 5.1602 ms/op 6.0717 ms/op 0.85
BeaconState.hashTreeRoot - 1 validator.effectiveBalance 68.608 us/op 63.773 us/op 1.08
BeaconState.hashTreeRoot - 32 validator.effectiveBalance 911.81 us/op 919.28 us/op 0.99
BeaconState.hashTreeRoot - 512 validator.effectiveBalance 12.795 ms/op 11.420 ms/op 1.12
BeaconState.hashTreeRoot - 1 balances 50.429 us/op 50.333 us/op 1.00
BeaconState.hashTreeRoot - 32 balances 476.36 us/op 446.61 us/op 1.07
BeaconState.hashTreeRoot - 512 balances 4.9372 ms/op 4.1590 ms/op 1.19
BeaconState.hashTreeRoot - 250000 balances 77.493 ms/op 72.379 ms/op 1.07
aggregationBits - 2048 els - zipIndexesInBitList 17.511 us/op 15.664 us/op 1.12
regular array get 100000 times 35.176 us/op 31.040 us/op 1.13
wrappedArray get 100000 times 33.968 us/op 40.040 us/op 0.85
arrayWithProxy get 100000 times 16.488 ms/op 14.990 ms/op 1.10
ssz.Root.equals 608.00 ns/op 571.00 ns/op 1.06
byteArrayEquals 586.00 ns/op 523.00 ns/op 1.12
shuffle list - 16384 els 7.2274 ms/op 6.6444 ms/op 1.09
shuffle list - 250000 els 105.48 ms/op 97.365 ms/op 1.08
processSlot - 1 slots 9.0310 us/op 8.1240 us/op 1.11
processSlot - 32 slots 1.4034 ms/op 1.3548 ms/op 1.04
getEffectiveBalanceIncrementsZeroInactive - 250000 vs - 7PWei 38.151 ms/op 36.849 ms/op 1.04
getCommitteeAssignments - req 1 vs - 250000 vc 3.2310 ms/op 2.8350 ms/op 1.14
getCommitteeAssignments - req 100 vs - 250000 vc 4.5366 ms/op 4.0190 ms/op 1.13
getCommitteeAssignments - req 1000 vs - 250000 vc 4.6790 ms/op 4.3471 ms/op 1.08
RootCache.getBlockRootAtSlot - 250000 vs - 7PWei 5.2500 ns/op 4.3500 ns/op 1.21
state getBlockRootAtSlot - 250000 vs - 7PWei 623.64 ns/op 980.28 ns/op 0.64
computeProposers - vc 250000 10.863 ms/op 10.590 ms/op 1.03
computeEpochShuffling - vc 250000 107.42 ms/op 101.64 ms/op 1.06
getNextSyncCommittee - vc 250000 190.16 ms/op 168.62 ms/op 1.13
computeSigningRoot for AttestationData 14.285 us/op 13.123 us/op 1.09
hash AttestationData serialized data then Buffer.toString(base64) 2.7385 us/op 2.3062 us/op 1.19
toHexString serialized data 1.2990 us/op 1.0478 us/op 1.24
Buffer.toString(base64) 397.94 ns/op 317.89 ns/op 1.25

by benchmarkbot/action

@twoeths

twoeths commented Apr 15, 2023

Copy link
Copy Markdown
Member Author

Skip deserializing attestation gossip messages turns out to be great, got better metrics than unstable (with same mesh peers, network i/o). Below is a feat2 mainnet metrics vs unstable mainnet metrics:

  • beacon_attestation gossip job time is 1/2 - 1/3 compared to unstable, the job wait time is reduced accordingly

Screenshot 2023-04-15 at 15 13 07

vs unstable
Screenshot 2023-04-15 at 15 13 53

  • notifyNewPayload time is 2x faster than unstable

Screenshot 2023-04-15 at 15 20 27

vs unstable
Screenshot 2023-04-15 at 15 21 08

  • "Gossip block process time" on feat2 (thanks to reduced notifyNewPayload)

Screenshot 2023-04-15 at 15 26 37

vs unstable
Screenshot 2023-04-15 at 15 27 07

  • "Gossip block received delay" on feat2

Screenshot 2023-04-15 at 15 35 49

vs unstable
Screenshot 2023-04-15 at 15 36 10

along with that:

  • Almost no dropped attestation gossip messages after 20h of testing
  • Better seen attestation data hit
  • Same heap space used, gc is a little bit better
  • Same attestation subnet mesh peers (which make it same gossip messages sent/received)

@twoeths twoeths changed the title Gossip validation for attestation: cache AttestationData Skip deserializing gossip attestation messages by caching AttestationData Apr 15, 2023
@twoeths
twoeths marked this pull request as ready for review April 15, 2023 08:48
@twoeths
twoeths requested a review from a team as a code owner April 15, 2023 08:48

@dapplion dapplion left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice optimization! It's unfortunate that deserializing is slow, let's monitor the RAM usage on a prod deployment. Please address the TODO on another PR when possible

* This is copied from ssz bitList.ts
* TODO: export this util from there
*/
function deserializeUint8ArrayBitListFromBytes(data: Uint8Array, start: number, end: number): BitArrayDeserialized {

@wemeetagain
wemeetagain merged commit fd00c1f into unstable Apr 17, 2023
@wemeetagain
wemeetagain deleted the tuyen/cache_attestation_data branch April 17, 2023 16:58
@wemeetagain

Copy link
Copy Markdown
Member

🎉 This PR is included in v1.8.0 🎉

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants