A Facebook discussion about Malta’s receipt lottery made me want to look at the data. Were there repeat winners across different draws? How much had they won? And were the winning receipts for small everyday purchases, larger amounts, or a mixture?
First, I needed to bring the data together.
I built a scraper to collect the publicly available PDFs from the VAT site, then used OCR to extract their contents into CSV files. From there, I started exploring the data in Jupyter notebooks.
I’m also building a site to present the results, taking inspiration from The Pudding’s approach to visual storytelling. I want to make the analysis something people can follow and explore.
The scraper runs daily through GitHub Actions to check for new source documents.
Next up is finishing the checks and publishing the results on the site.