Fast Catchup not working for our node on MainNet

My node is also failing to succeed with the Fast Catchup on mainnet.

After 6 hours it didn’t seem to be getting any further. The syptoms visible in the status were that the sync time continued to climb, the “Catchpoint accounts processed” climbed for a while and then started back at the beginning to climb a bit higher the next time. After doing this several times it seemed unable to get passed 1113600 processed catchpoint accounts.

$ goal node status -d /var/lib/algorand -w 1000
Last committed block: 8570
Sync Time: 2382.1s
Catchpoint: 8560000#JXTHWW5BRW7MDSPVHITBLYORCNUDTHIXZRIYOZYKHDHGU2WUXLZQ
Catchpoint total accounts: 4313466
Catchpoint accounts processed: 1048064
Genesis ID: mainnet-v1.0
Genesis hash: wGHE2Pwdvd7S12BL5FaOP20EGYesN73ktiC1qzkkit8=
^C

$ goal node status -d /var/lib/algorand -w 1000
Last committed block: 8570
Sync Time: 22162.2s
Catchpoint: 8560000#JXTHWW5BRW7MDSPVHITBLYORCNUDTHIXZRIYOZYKHDHGU2WUXLZQ
Catchpoint total accounts: 4313466
Catchpoint accounts processed: 1113600
Genesis ID: mainnet-v1.0
Genesis hash: wGHE2Pwdvd7S12BL5FaOP20EGYesN73ktiC1qzkkit8=
^C

I gave up and aborted the catchup and stopped algod.

A quick look in node.log showed repeated messages along the lines of:

{“callee”:“github.com/algorand/go-algorand/ledger.(*CatchpointCatchupAccessorImpl) .processStagingBalances.func1”,“caller”:“/root/go/s
rc/github.com/algorand/go-algorand/ledger/catchupaccessor.go:294”,“file”:“dbutil.go”,“function”:“github.com/algorand/go-algorand/util
/db.(*Accessor).atomic”,“level”:“warning”,“line”:355,“msg”:“dbatomic: tx surpassed expected deadline by 1.131643661s”,“name”:“”,“read
only”:false,“time”:“2020-08-20T11:12:38.108569+01:00”}
{“file”:“catchpointService.go”,“function”:“github.com/algorand/go-algorand/catch
up.(*CatchpointCatchupService).processStageLedgerDownload”,“level”:“warning”,“li
ne”:293,“msg”:“unable to download ledger : unexpected EOF”,“name”:“”,“time”:“202
0-08-20T12:22:31.222640+01:00”}

I see @cusma has the same symptoms.

From what @jimmy has posted it sounds as though there may be little hope for my ancient node untl I give it a memory upgrade beyond 4GB.

I’ll try increasing the CatchupBlockDownloadRetryAttempts - perhaps to 8000 and let you know how it goes.

I’ve not yet found any documentation for recomended node specifications - can you point me to it?