Parsers and data collection for SEO and marketing
Parsers and data collection are part of the Process automation service at Best Boost Partners. We write parsers that gather search results, competitor prices and catalogues, backlink data and page metadata on a schedule, clean the data and store it in a database or spreadsheet. Where the source has an official API we work through it, and we follow the source site’s terms and robots rules. You receive the code, the documentation and the first collected dataset.
What kinds of parsers do we write?
Search results
We collect the search results for your list of keywords by region and device: positions, addresses, titles, snippets and the special blocks on the page. The data shows who actually ranks for each query and how the results change over time.
Competitor prices and catalogues
The parser goes through competitor categories and product pages and records names, prices, availability and attributes. An online store or an auto parts catalogue gets a comparison table that is refreshed on a schedule.
Backlink data
We take backlink data through the APIs of link indexes such as Ahrefs and Majestic, under your subscription, and check the live pages: whether the link is in place, which anchor and attributes it has, which status code the page returns.
Page metadata
A crawler reads titles, descriptions, headings, canonical tags, status codes and structured data from your pages or from competitor pages. The result goes into a table that can be filtered and compared with the previous run.
Cleaning and storage
Raw data is cleared of duplicates, brought to one format and checked before it is saved. We store it in a database or in a spreadsheet: the choice depends on the volume and on who will work with the data.
Schedule and control
Each parser runs on a schedule and keeps a log. If a run returns an empty or unusually small result, the parser sends an alert: this usually means the source has changed its layout.
What do you get from parsers and data collection?
- Parser code in your repository, with the settings kept in a separate configuration file.
- A database or spreadsheet with a described structure: what each field means and where it comes from.
- A schedule of runs and a log of every run.
- Documentation on how to add a new source, a keyword list or a field.
- A note on each source: API or parsing, request limits, the terms of use that were taken into account.
When are parsers and data collection needed?
- The same data is regularly copied by hand from search results or from competitor sites.
- An online store needs to see competitor prices and availability every day, and the catalogue is too large to check by hand.
- Bought or placed links have to be checked regularly: whether the page is live and the link is still in place.
- A decision on keywords or site structure needs data on many competitor pages at once.
Frequently asked questions
Is it permitted to parse someone else’s site?
It depends on the source site’s terms of use, the type of data and the law of the country, so we check every source before the work starts. We use the official API where there is one, follow robots rules, keep the request rate moderate and do not collect personal data or content closed behind a login. If a source forbids automated collection, we say so and look for another source of the same data.
What happens if the source site changes its layout?
The parser stops finding the fields it needs, records the error and sends an alert. Selectors and field rules are kept in the configuration, apart from the main code, so the correction is usually small. The documentation describes how to make the correction.
How often can the data be collected?
As often as the task requires and the source allows. Prices and availability are usually collected daily, page metadata once a week or after a release, link checks within the limits of the index API. We set the frequency for each source separately and write it into the documentation.
Tell us about the project
Fill in what you already know. Blank fields are fine: we will clarify the rest in conversation.
Brief sent
Thank you. We will reply within one working day.
Preview mode: the form handler is not connected yet, so nothing was sent.