Scrape Genie users — is this normal, or have I got something badly configured?
I'm trying to scrape some email leads and the performance seems incredibly slow.
My current setup
Search:
- 7 keywords
- 43 UK locations/cities/areas
- 5 search engines:
- 4get Ghostery GB
- 4get Yep GB
- Bing GB
- DuckDuckGo UK
- Google GB
- I originally had Google USA and DuckDuckGo USA enabled, but removed them because they were producing huge numbers of irrelevant US results. Even with the selected cities I was using.
- Start page: 1
- Pages to parse: 100
- Extract only root domains: ON
- Use proxies for search engines: ON
- Skip keywords already processed: ON
- Repeat search when finished: OFF
- Add cities: ON
Parser:
- Wait after download: OFF
- Use proxies for parsing: OFF
- Accept duplicate domains: OFF
- Parse for sublinks: ON
- Sublinks: 1 level deep
- Currently set to within root-domain
- Follow sublinks only if URL filter matches: ON
- I am trying to use this to find emails on contact/about pages rather than just the landing page.
- I am extracting emails only — I'm not interested in phone numbers or addresses.
Global settings:
- Checking/submission threads: 100
- Scraping/search threads: 100
- HTTP timeout: 60 seconds
- Automatic delay between search queries
- Stop projects when no proxy is alive: ON
Proxies:
- 25 proxies
- All 25 successfully tested (Static Residential (ISP))
- Proxy test showed 25/25 working
- Proxy types: Connect, SOCKS4 and SOCKS5
- Web proxies aren't enabled
- Keep-alive: ON
- Randomise list / avoid fake positive port scans: ON
- Proxy timeout: 5 seconds
- 20 threads in the proxy testing/configuration section
What I'm actually getting
My previous runs were:
10 threads: about 352 scraped email entries after more than a day.
50 threads: about 1,091 entries after roughly a day.
I then increased to 100 scraping + 100 parsing threads.
The current run has now been going for 23 hours 29 minutes and has produced 1,273 scraped email entries.
It has used about 6.6 GB of proxy traffic in that time.
So I'm getting roughly 50–55 email entries per hour.
And this is the bit I don't understand.
Going from 10 → 50 threads made a huge difference.
Going from 50 → 100 threads has made relatively little difference.
I have 25 working proxies, but 100 scraping threads and 100 parsing threads.
There is another thing worrying me
Looking at the parser log, it seems to be doing a LOT of work on individual websites.
I've seen it parsing things such as:
- CSS
- JavaScript
- PNG/JPG/WebP images
- WordPress plugin files
/wp-json/- AJAX endpoints
- PDFs
- various other technical URLs
rather than simply:
search result → website → contact/about page → email → next website
I'm wondering whether the one-level sublink parsing is causing it to crawl far more of each website than I intended.
So what am I doing wrong?
Is this performance normal for Scrape Genie?
Or is there something fundamentally wrong with my configuration?
In particular:
Does 25 proxies simply not give enough capacity for 100 scraping + 100 parsing threads?
Do people running Scrape Genie at a useful commercial lead-generation speed really need something like 50–100+ proxies?
Or should 25 proxies be perfectly capable of producing substantially more than ~1,300 scraped email entries per day?
I'm not expecting 100 threads to magically mean 100× the speed. I just don't understand why, once I went from 50 to 100 threads, the throughput barely increased.
I'm also wondering whether my 1-level sublink parsing is creating a massive amount of unnecessary traffic because it's crawling assets and technical URLs.
I'd really appreciate someone who knows Scrape Genie giving me an idea of where the bottleneck is likely to be and what settings I should actually be looking at.
I'd love it to be working a lot quicker.
Because at the moment, I could probably scrape the same or more leads myself in 24 hours with just the one browser and a lot of coffee.
Comments