yes. the problem is, we are fetching products from an API. and since we thought processing power will be a limiting factor, we thought that sorting out duplicates would reduce load.
but since the different microservices which process the data are taking different times we are using the sql tables as a pool. this should help upcscaling by using multiple microservices.
cloud services are yet not a solution as we are still in development.
I need to dedupe the to-be-processed data with the data thats already processed in the "final" table. We are working with hundreds of millions of products therefore we thought about "simply" using random batches from the data to be processed. But thanks to the many replies Ive learned already that our approach was in the beginning already wrong.