Skip to content

Limit preaggregating attestations - #5256

Merged
wemeetagain merged 6 commits into
unstablefrom
tuyen/attestationPool
Mar 29, 2023
Merged

Limit preaggregating attestations#5256
wemeetagain merged 6 commits into
unstablefrom
tuyen/attestationPool

Conversation

@twoeths

@twoeths twoeths commented Mar 10, 2023

Copy link
Copy Markdown
Member

Motivation

As shown on a subscribe-all-subnets goerli node, preaggregating attestations takes 8% of cpu time, on mainnet it could be worse. I think some preaggregation are redundant because validators only get AggregatedAttestation at 2/3 of clock slot, including:

  • Attestations with slot < clockSlot and > clockSlot - 3
  • Attestations with slot = clockSlot but come to the pool at > 2/3 of slot

I noticed that on mainnet node, the attestation job wait time is 7s in average at some points so they are not useful to preaggregate anymore

Description

  • Do not preaggregate attestations from old slots
  • Do no preaggregate attestations in clock slot but come to the pool at > 2/3 of slot
  • Same to SyncCommitteeMessage

part of #5247

@github-actions

github-actions Bot commented Mar 10, 2023

Copy link
Copy Markdown
Contributor

Performance Report

鉁旓笍 no performance regression detected

Full benchmark results
Benchmark suite Current: 3317fc8 Previous: b861ab8 Ratio
getPubkeys - index2pubkey - req 1000 vs - 250000 vc 725.63 us/op 902.87 us/op 0.80
getPubkeys - validatorsArr - req 1000 vs - 250000 vc 48.291 us/op 46.746 us/op 1.03
BLS verify - blst-native 1.2178 ms/op 1.2268 ms/op 0.99
BLS verifyMultipleSignatures 3 - blst-native 2.4904 ms/op 2.4862 ms/op 1.00
BLS verifyMultipleSignatures 8 - blst-native 5.3441 ms/op 5.3377 ms/op 1.00
BLS verifyMultipleSignatures 32 - blst-native 19.106 ms/op 19.364 ms/op 0.99
BLS aggregatePubkeys 32 - blst-native 25.866 us/op 26.459 us/op 0.98
BLS aggregatePubkeys 128 - blst-native 100.88 us/op 101.51 us/op 0.99
getAttestationsForBlock 57.860 ms/op 52.723 ms/op 1.10
isKnown best case - 1 super set check 261.00 ns/op 246.00 ns/op 1.06
isKnown normal case - 2 super set checks 253.00 ns/op 240.00 ns/op 1.05
isKnown worse case - 16 super set checks 249.00 ns/op 244.00 ns/op 1.02
CheckpointStateCache - add get delete 4.9870 us/op 4.8540 us/op 1.03
validate gossip signedAggregateAndProof - struct 2.7547 ms/op 2.7776 ms/op 0.99
validate gossip attestation - struct 1.3245 ms/op 1.3290 ms/op 1.00
pickEth1Vote - no votes 1.3462 ms/op 1.2548 ms/op 1.07
pickEth1Vote - max votes 9.7029 ms/op 11.555 ms/op 0.84
pickEth1Vote - Eth1Data hashTreeRoot value x2048 9.2195 ms/op 9.1310 ms/op 1.01
pickEth1Vote - Eth1Data hashTreeRoot tree x2048 15.243 ms/op 14.489 ms/op 1.05
pickEth1Vote - Eth1Data fastSerialize value x2048 658.73 us/op 648.69 us/op 1.02
pickEth1Vote - Eth1Data fastSerialize tree x2048 8.0580 ms/op 7.7747 ms/op 1.04
bytes32 toHexString 485.00 ns/op 478.00 ns/op 1.01
bytes32 Buffer.toString(hex) 365.00 ns/op 343.00 ns/op 1.06
bytes32 Buffer.toString(hex) from Uint8Array 564.00 ns/op 531.00 ns/op 1.06
bytes32 Buffer.toString(hex) + 0x 348.00 ns/op 333.00 ns/op 1.05
Object access 1 prop 0.15800 ns/op 0.16400 ns/op 0.96
Map access 1 prop 0.15800 ns/op 0.15600 ns/op 1.01
Object get x1000 6.7970 ns/op 6.4860 ns/op 1.05
Map get x1000 0.59900 ns/op 0.56000 ns/op 1.07
Object set x1000 55.933 ns/op 49.833 ns/op 1.12
Map set x1000 44.244 ns/op 41.887 ns/op 1.06
Return object 10000 times 0.23740 ns/op 0.22860 ns/op 1.04
Throw Error 10000 times 4.1619 us/op 4.0222 us/op 1.03
fastMsgIdFn sha256 / 200 bytes 3.4320 us/op 3.3280 us/op 1.03
fastMsgIdFn h32 xxhash / 200 bytes 274.00 ns/op 270.00 ns/op 1.01
fastMsgIdFn h64 xxhash / 200 bytes 400.00 ns/op 378.00 ns/op 1.06
fastMsgIdFn sha256 / 1000 bytes 11.788 us/op 11.310 us/op 1.04
fastMsgIdFn h32 xxhash / 1000 bytes 408.00 ns/op 396.00 ns/op 1.03
fastMsgIdFn h64 xxhash / 1000 bytes 473.00 ns/op 448.00 ns/op 1.06
fastMsgIdFn sha256 / 10000 bytes 103.95 us/op 100.95 us/op 1.03
fastMsgIdFn h32 xxhash / 10000 bytes 1.9100 us/op 1.8620 us/op 1.03
fastMsgIdFn h64 xxhash / 10000 bytes 1.3710 us/op 1.3250 us/op 1.03
enrSubnets - fastDeserialize 64 bits 1.2780 us/op 1.2530 us/op 1.02
enrSubnets - ssz BitVector 64 bits 489.00 ns/op 462.00 ns/op 1.06
enrSubnets - fastDeserialize 4 bits 170.00 ns/op 165.00 ns/op 1.03
enrSubnets - ssz BitVector 4 bits 490.00 ns/op 468.00 ns/op 1.05
prioritizePeers score -10:0 att 32-0.1 sync 2-0 101.95 us/op 104.63 us/op 0.97
prioritizePeers score 0:0 att 32-0.25 sync 2-0.25 131.53 us/op 129.97 us/op 1.01
prioritizePeers score 0:0 att 32-0.5 sync 2-0.5 169.38 us/op 171.67 us/op 0.99
prioritizePeers score 0:0 att 64-0.75 sync 4-0.75 300.01 us/op 302.50 us/op 0.99
prioritizePeers score 0:0 att 64-1 sync 4-1 363.48 us/op 364.82 us/op 1.00
array of 16000 items push then shift 1.6514 us/op 1.6274 us/op 1.01
LinkedList of 16000 items push then shift 8.8550 ns/op 8.8590 ns/op 1.00
array of 16000 items push then pop 94.137 ns/op 78.641 ns/op 1.20
LinkedList of 16000 items push then pop 8.6320 ns/op 8.6300 ns/op 1.00
array of 24000 items push then shift 2.3456 us/op 2.3913 us/op 0.98
LinkedList of 24000 items push then shift 8.7640 ns/op 8.8530 ns/op 0.99
array of 24000 items push then pop 86.534 ns/op 74.127 ns/op 1.17
LinkedList of 24000 items push then pop 8.5440 ns/op 8.3720 ns/op 1.02
intersect bitArray bitLen 8 13.422 ns/op 13.167 ns/op 1.02
intersect array and set length 8 79.146 ns/op 76.914 ns/op 1.03
intersect bitArray bitLen 128 44.303 ns/op 43.571 ns/op 1.02
intersect array and set length 128 1.1004 us/op 1.0424 us/op 1.06
Buffer.concat 32 items 2.6970 us/op 2.5760 us/op 1.05
Uint8Array.set 32 items 2.6000 us/op 2.9030 us/op 0.90
pass gossip attestations to forkchoice per slot 3.2684 ms/op 3.1369 ms/op 1.04
computeDeltas 2.8473 ms/op 3.3523 ms/op 0.85
computeProposerBoostScoreFromBalances 1.7781 ms/op 1.7936 ms/op 0.99
altair processAttestation - 250000 vs - 7PWei normalcase 2.2409 ms/op 2.1321 ms/op 1.05
altair processAttestation - 250000 vs - 7PWei worstcase 3.5585 ms/op 3.3587 ms/op 1.06
altair processAttestation - setStatus - 1/6 committees join 141.18 us/op 142.07 us/op 0.99
altair processAttestation - setStatus - 1/3 committees join 275.35 us/op 279.48 us/op 0.99
altair processAttestation - setStatus - 1/2 committees join 373.86 us/op 358.46 us/op 1.04
altair processAttestation - setStatus - 2/3 committees join 461.15 us/op 446.96 us/op 1.03
altair processAttestation - setStatus - 4/5 committees join 648.38 us/op 636.80 us/op 1.02
altair processAttestation - setStatus - 100% committees join 755.24 us/op 752.99 us/op 1.00
altair processBlock - 250000 vs - 7PWei normalcase 16.033 ms/op 18.435 ms/op 0.87
altair processBlock - 250000 vs - 7PWei normalcase hashState 25.894 ms/op 28.389 ms/op 0.91
altair processBlock - 250000 vs - 7PWei worstcase 51.566 ms/op 47.485 ms/op 1.09
altair processBlock - 250000 vs - 7PWei worstcase hashState 66.906 ms/op 65.453 ms/op 1.02
phase0 processBlock - 250000 vs - 7PWei normalcase 1.9865 ms/op 1.9827 ms/op 1.00
phase0 processBlock - 250000 vs - 7PWei worstcase 29.750 ms/op 26.991 ms/op 1.10
altair processEth1Data - 250000 vs - 7PWei normalcase 490.69 us/op 451.59 us/op 1.09
vc - 250000 eb 1 eth1 1 we 0 wn 0 - smpl 15 8.7450 us/op 6.7240 us/op 1.30
vc - 250000 eb 0.95 eth1 0.1 we 0.05 wn 0 - smpl 219 26.656 us/op 19.591 us/op 1.36
vc - 250000 eb 0.95 eth1 0.3 we 0.05 wn 0 - smpl 42 11.238 us/op 8.2370 us/op 1.36
vc - 250000 eb 0.95 eth1 0.7 we 0.05 wn 0 - smpl 18 8.2570 us/op 6.1490 us/op 1.34
vc - 250000 eb 0.1 eth1 0.1 we 0 wn 0 - smpl 1020 91.798 us/op 74.511 us/op 1.23
vc - 250000 eb 0.03 eth1 0.03 we 0 wn 0 - smpl 11777 642.71 us/op 627.01 us/op 1.03
vc - 250000 eb 0.01 eth1 0.01 we 0 wn 0 - smpl 16384 917.38 us/op 910.46 us/op 1.01
vc - 250000 eb 0 eth1 0 we 0 wn 0 - smpl 16384 878.83 us/op 854.80 us/op 1.03
vc - 250000 eb 0 eth1 0 we 0 wn 0 nocache - smpl 16384 2.3693 ms/op 2.2624 ms/op 1.05
vc - 250000 eb 0 eth1 1 we 0 wn 0 - smpl 16384 1.7461 ms/op 1.4799 ms/op 1.18
vc - 250000 eb 0 eth1 1 we 0 wn 0 nocache - smpl 16384 3.9584 ms/op 3.7144 ms/op 1.07
Tree 40 250000 create 318.68 ms/op 294.90 ms/op 1.08
Tree 40 250000 get(125000) 182.07 ns/op 177.52 ns/op 1.03
Tree 40 250000 set(125000) 875.31 ns/op 841.54 ns/op 1.04
Tree 40 250000 toArray() 17.529 ms/op 16.070 ms/op 1.09
Tree 40 250000 iterate all - toArray() + loop 16.673 ms/op 16.193 ms/op 1.03
Tree 40 250000 iterate all - get(i) 68.708 ms/op 63.605 ms/op 1.08
MutableVector 250000 create 9.4076 ms/op 9.9340 ms/op 0.95
MutableVector 250000 get(125000) 6.3260 ns/op 6.0540 ns/op 1.04
MutableVector 250000 set(125000) 272.26 ns/op 248.56 ns/op 1.10
MutableVector 250000 toArray() 2.8788 ms/op 2.7310 ms/op 1.05
MutableVector 250000 iterate all - toArray() + loop 2.9921 ms/op 2.8652 ms/op 1.04
MutableVector 250000 iterate all - get(i) 1.5340 ms/op 1.4489 ms/op 1.06
Array 250000 create 2.6896 ms/op 2.4904 ms/op 1.08
Array 250000 clone - spread 1.2721 ms/op 1.2991 ms/op 0.98
Array 250000 get(125000) 0.61700 ns/op 0.60300 ns/op 1.02
Array 250000 set(125000) 0.69800 ns/op 0.67100 ns/op 1.04
Array 250000 iterate all - loop 83.571 us/op 82.726 us/op 1.01
effectiveBalanceIncrements clone Uint8Array 300000 27.431 us/op 27.322 us/op 1.00
effectiveBalanceIncrements clone MutableVector 300000 408.00 ns/op 404.00 ns/op 1.01
effectiveBalanceIncrements rw all Uint8Array 300000 166.61 us/op 168.43 us/op 0.99
effectiveBalanceIncrements rw all MutableVector 300000 84.405 ms/op 82.996 ms/op 1.02
phase0 afterProcessEpoch - 250000 vs - 7PWei 113.08 ms/op 113.59 ms/op 1.00
phase0 beforeProcessEpoch - 250000 vs - 7PWei 34.975 ms/op 42.867 ms/op 0.82
altair processEpoch - mainnet_e81889 329.82 ms/op 305.02 ms/op 1.08
mainnet_e81889 - altair beforeProcessEpoch 63.016 ms/op 63.387 ms/op 0.99
mainnet_e81889 - altair processJustificationAndFinalization 16.169 us/op 17.679 us/op 0.91
mainnet_e81889 - altair processInactivityUpdates 5.0869 ms/op 5.3227 ms/op 0.96
mainnet_e81889 - altair processRewardsAndPenalties 68.260 ms/op 53.641 ms/op 1.27
mainnet_e81889 - altair processRegistryUpdates 2.7230 us/op 2.6770 us/op 1.02
mainnet_e81889 - altair processSlashings 536.00 ns/op 612.00 ns/op 0.88
mainnet_e81889 - altair processEth1DataReset 468.00 ns/op 786.00 ns/op 0.60
mainnet_e81889 - altair processEffectiveBalanceUpdates 1.2275 ms/op 1.3085 ms/op 0.94
mainnet_e81889 - altair processSlashingsReset 4.6350 us/op 5.5330 us/op 0.84
mainnet_e81889 - altair processRandaoMixesReset 4.4010 us/op 4.8890 us/op 0.90
mainnet_e81889 - altair processHistoricalRootsUpdate 904.00 ns/op 665.00 ns/op 1.36
mainnet_e81889 - altair processParticipationFlagUpdates 2.0750 us/op 2.3090 us/op 0.90
mainnet_e81889 - altair processSyncCommitteeUpdates 612.00 ns/op 612.00 ns/op 1.00
mainnet_e81889 - altair afterProcessEpoch 124.13 ms/op 123.92 ms/op 1.00
phase0 processEpoch - mainnet_e58758 364.07 ms/op 319.33 ms/op 1.14
mainnet_e58758 - phase0 beforeProcessEpoch 133.16 ms/op 124.57 ms/op 1.07
mainnet_e58758 - phase0 processJustificationAndFinalization 16.866 us/op 15.435 us/op 1.09
mainnet_e58758 - phase0 processRewardsAndPenalties 64.169 ms/op 55.397 ms/op 1.16
mainnet_e58758 - phase0 processRegistryUpdates 7.1590 us/op 8.9250 us/op 0.80
mainnet_e58758 - phase0 processSlashings 535.00 ns/op 535.00 ns/op 1.00
mainnet_e58758 - phase0 processEth1DataReset 483.00 ns/op 599.00 ns/op 0.81
mainnet_e58758 - phase0 processEffectiveBalanceUpdates 967.37 us/op 994.36 us/op 0.97
mainnet_e58758 - phase0 processSlashingsReset 4.2910 us/op 3.3820 us/op 1.27
mainnet_e58758 - phase0 processRandaoMixesReset 4.8570 us/op 4.5490 us/op 1.07
mainnet_e58758 - phase0 processHistoricalRootsUpdate 525.00 ns/op 883.00 ns/op 0.59
mainnet_e58758 - phase0 processParticipationRecordUpdates 3.9150 us/op 5.6430 us/op 0.69
mainnet_e58758 - phase0 afterProcessEpoch 96.691 ms/op 96.639 ms/op 1.00
phase0 processEffectiveBalanceUpdates - 250000 normalcase 1.2425 ms/op 1.2713 ms/op 0.98
phase0 processEffectiveBalanceUpdates - 250000 worstcase 0.5 1.4822 ms/op 1.5035 ms/op 0.99
altair processInactivityUpdates - 250000 normalcase 25.490 ms/op 26.604 ms/op 0.96
altair processInactivityUpdates - 250000 worstcase 25.953 ms/op 37.941 ms/op 0.68
phase0 processRegistryUpdates - 250000 normalcase 6.6910 us/op 20.655 us/op 0.32
phase0 processRegistryUpdates - 250000 badcase_full_deposits 231.19 us/op 538.22 us/op 0.43
phase0 processRegistryUpdates - 250000 worstcase 0.5 123.28 ms/op 208.78 ms/op 0.59
altair processRewardsAndPenalties - 250000 normalcase 68.964 ms/op 100.51 ms/op 0.69
altair processRewardsAndPenalties - 250000 worstcase 69.482 ms/op 105.94 ms/op 0.66
phase0 getAttestationDeltas - 250000 normalcase 6.2917 ms/op 12.856 ms/op 0.49
phase0 getAttestationDeltas - 250000 worstcase 6.4882 ms/op 13.070 ms/op 0.50
phase0 processSlashings - 250000 worstcase 3.4779 ms/op 7.5046 ms/op 0.46
altair processSyncCommitteeUpdates - 250000 172.14 ms/op 220.84 ms/op 0.78
BeaconState.hashTreeRoot - No change 260.00 ns/op 354.00 ns/op 0.73
BeaconState.hashTreeRoot - 1 full validator 53.142 us/op 71.761 us/op 0.74
BeaconState.hashTreeRoot - 32 full validator 481.87 us/op 678.42 us/op 0.71
BeaconState.hashTreeRoot - 512 full validator 5.5427 ms/op 7.1447 ms/op 0.78
BeaconState.hashTreeRoot - 1 validator.effectiveBalance 61.594 us/op 66.046 us/op 0.93
BeaconState.hashTreeRoot - 32 validator.effectiveBalance 888.91 us/op 1.0266 ms/op 0.87
BeaconState.hashTreeRoot - 512 validator.effectiveBalance 11.744 ms/op 14.820 ms/op 0.79
BeaconState.hashTreeRoot - 1 balances 48.434 us/op 54.563 us/op 0.89
BeaconState.hashTreeRoot - 32 balances 449.58 us/op 580.07 us/op 0.78
BeaconState.hashTreeRoot - 512 balances 4.2744 ms/op 5.9385 ms/op 0.72
BeaconState.hashTreeRoot - 250000 balances 72.568 ms/op 105.57 ms/op 0.69
aggregationBits - 2048 els - zipIndexesInBitList 15.197 us/op 39.111 us/op 0.39
regular array get 100000 times 32.709 us/op 41.625 us/op 0.79
wrappedArray get 100000 times 32.727 us/op 36.893 us/op 0.89
arrayWithProxy get 100000 times 15.960 ms/op 16.780 ms/op 0.95
ssz.Root.equals 542.00 ns/op 720.00 ns/op 0.75
byteArrayEquals 532.00 ns/op 810.00 ns/op 0.66
shuffle list - 16384 els 6.7622 ms/op 7.7617 ms/op 0.87
shuffle list - 250000 els 99.037 ms/op 126.50 ms/op 0.78
processSlot - 1 slots 8.4610 us/op 13.581 us/op 0.62
processSlot - 32 slots 1.3201 ms/op 1.6638 ms/op 0.79
getEffectiveBalanceIncrementsZeroInactive - 250000 vs - 7PWei 38.078 ms/op 39.744 ms/op 0.96
getCommitteeAssignments - req 1 vs - 250000 vc 2.8756 ms/op 2.9538 ms/op 0.97
getCommitteeAssignments - req 100 vs - 250000 vc 4.0789 ms/op 4.1995 ms/op 0.97
getCommitteeAssignments - req 1000 vs - 250000 vc 4.4335 ms/op 4.4752 ms/op 0.99
RootCache.getBlockRootAtSlot - 250000 vs - 7PWei 4.6100 ns/op 4.7400 ns/op 0.97
state getBlockRootAtSlot - 250000 vs - 7PWei 972.20 ns/op 1.0465 us/op 0.93
computeProposers - vc 250000 10.369 ms/op 12.489 ms/op 0.83
computeEpochShuffling - vc 250000 100.04 ms/op 111.58 ms/op 0.90
getNextSyncCommittee - vc 250000 172.19 ms/op 200.11 ms/op 0.86

by benchmarkbot/action

@twoeths
twoeths marked this pull request as ready for review March 10, 2023 10:05
@twoeths
twoeths requested a review from a team as a code owner March 10, 2023 10:05

@dapplion dapplion left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure about the safety of this change in adverse conditions. This settings should be flags in case they need to be customized in the future

@twoeths

twoeths commented Mar 12, 2023

Copy link
Copy Markdown
Member Author

@dapplion are you worry about the lowestPemissibleSlot condition or cutOffSecFromSlot condition or both?

  • I don't see it's reasonable to preaggregate attestations of old slots as aggregators only need to aggregate attestations of current slot
  • For cutOffSecFromSlot condition we could accept toleranceSec which is 0.5s => do you mean a flag for this?

@dapplion

Copy link
Copy Markdown
Contributor

Only lowestPemissibleSlot, default to 1 (only same slot) but allow to increase to more

@twoeths
twoeths force-pushed the tuyen/attestationPool branch from 23b11ed to bdefdef Compare March 13, 2023 23:07

// validator gets SyncCommitteeContribution at 2/3 of slot, it's no use to preaggregate later than that time
if (this.clock.secFromSlot(slot) > this.cutOffSecFromSlot) {
throw new OpPoolError({code: OpPoolErrorCode.LATE_MESSAGE, slot});

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

How loud would be this error?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we throw error here and gossipHandler would log error there.
since this is mainly for debugging purpose, not for end user so I changed log level to debug.

@dapplion dapplion Mar 19, 2023

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why not return an insert outcome? Then the caller can decide to maybe log a debug if result is not add

Comment thread packages/beacon-node/src/chain/opPools/attestationPool.ts Outdated
dapplion
dapplion previously approved these changes Mar 28, 2023
@wemeetagain
wemeetagain merged commit 2b03d90 into unstable Mar 29, 2023
@wemeetagain
wemeetagain deleted the tuyen/attestationPool branch March 29, 2023 03:26
twoeths added a commit that referenced this pull request Mar 29, 2023
* Limit preaggregating attestations

* preaggregateSlotDistance hidden cli param

* Log debug if error adding SyncCommitteeMessage to pool

* SyncCommitteeMessagePool: return instead of throw error

* Add SyncCommitteeMesssage insertOutcome metric

* Update prune() method header

Co-authored-by: Cayman <caymannava@gmail.com>

---------

Co-authored-by: Cayman <caymannava@gmail.com>
@wemeetagain

Copy link
Copy Markdown
Member

馃帀 This PR is included in v1.8.0 馃帀

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants