Making my own version of firecrawl when scraping language schools
A mate was hand-typing every German course in Berlin into one comparable spreadsheet, and it was taking a real toll on his sanity. So I built a chain that finds the schools, reads their sites, and pulls every class into one table.
This one wasn’t even my problem. A mate had it, I’d just noticed relevance.ai could search Google and open web pages, and I got curious whether I could build a scraper faster than he could get the data himself.
Context
He’d reached out with a great expat public service: aggregate all the language school data in Berlin into one comparable database, so you could actually find the best school for you. The dream being he could then do it for any language in any city in the world.
Sounds easy. There are 38 language schools in Berlin, each with tens of different class types, scattered randomly around their websites and expressed in very creative ways.
To his credit he’d gone through page after page of it, interpreted each offer so it could sit next to the others, and logged it all into a spreadsheet by hand. That’s where the sanity was going.
How it’s built
Four steps. The first three are chained in relevance.ai, then a Jupyter notebook does the tidying up.
- Perplexity
1. Discover schools
Every German language school in Berlin and its homepage.
- Relevance AI
2. Locate the offer pages
Pulls the HTML off each homepage and picks the page most likely to hold the class details.
- Claude
3. Extract course details
Reads each offer page, interprets what’s being sold, and writes it into a structured table.
- Jupyter
4. Flatten and organise
Python in a notebook. One table, so classes can be compared across schools.
The stack
- Perplexity · Finds the schools and their homepages. The most accurate model I tried at live information.
- relevance.ai · Runs the chain. Pulls the HTML off each homepage and picks the page with the class offers.
- Claude · Reads the offer pages and extracts the class details. Sonnet.
- Jupyter · A Python notebook that flattens the output into one table.
Step 3 is where all the effort went
Doing it all in one go didn’t work. I kept breaking the task down until each step was small enough for a model to get right, and step 3 still took the most trial and error by a mile. The prompt has to spell out what every column means and what the output should look like, or Claude misses context, fills the wrong column, and hands back something that won’t sit next to the next school.
The model was great at extracting the data out of raw HTML. The hard part was the work before it, cleaning the page down and structuring what was left so it fit inside a prompt. Once that was working, the crawling and scraping was general to any school.
Learnings
- Every scraper I’d built before only worked on structured sites, where I knew what the data looked like, where it sat, and how to move around to get more. This was two or three prompts reading a page cold. Given how random these sites are, that impressed me more than anything else in the build.
- I learnt about Firecrawl later, which leverages AI in the same way. We were solving the same problem at around the same time they were being founded, and that made me feel very cutting edge.
- No single model could do it. I ended up mixing Perplexity, an OpenAI model and Claude Sonnet across the chain, because each was clearly better at its own bit and clearly worse at the others.
- I didn’t just waste my life. I couldn’t ask Perplexity which school to go to and be done: it told me GLS, which is way too expensive for an unemployed man. The school I picked off the aggregated data is not the one I was recommended.
- The numbers came out close but not aligned. Close enough to spot the outliers, nowhere near close enough to sort a column by price per hour.
For a first 80/20 pass at a landscape, it’s a huge timesaver. As a scraper, it’s not the best in the world.
Next steps
What’s still broken
- It isn’t repeatable. Run it again and you get a different set of results, which is the one reason I’d never put it in production.
- It isn’t complete. It found the 24 most active schools out of the 38 in Berlin.
What I want to build
- Nothing else on this one. I’m done with this experiment, it’s taken enough of my life.
- The workflow is general to any language in any city, which was my friend’s actual dream, so that part is sitting there ready to point somewhere else.
This was exploratory. I ran it, picked a school off the output, and nobody ran it again. Maybe my sanity’s at the same level as my friend’s now.




