Data quality lessons
We found 5,951 jobs posted by a company called "CHECKED." It doesn't exist.
What a quarter-million job postings taught us about data quality
🚩 Placeholder values hide in plain sight
Values like "Checked" or "Unknown" sat in 5,951 and 1,230 postings respectively. Location/industry strings leaked into the company-name field too. Left unfiltered, a placeholder outranked most real employers.
🪞 The same company, many names
Legal-suffix variants and casing differences fragment one entity into several. Only ~12% of our distinct advertiser text values had ever been resolved into one clean record.
📎 Posting counts ≠ job counts
16,341 posting records sat inside same-day syndication bursts — one role alone showed 59 "postings" within 25 seconds. Raw counts overstate real vacancy numbers.
🕳️ Missing fields aren't random
Salary, postcode, and industry fields are missing at very different rates by source site — a metric that silently drops missing rows can quietly bias what you conclude.
🎯 Why this matters for your reporting
If your office publishes statistics to students, leadership, or accreditation bodies based on job market data — "X% of roles now list salary," "our top placement partners are Y and Z" — these are exactly the quiet distortions that can make a well-intentioned number misleading. Worth a data-quality pass before publication, not after someone asks a hard question about it.
Build on clean data
We do this cleanup work systematically at GetJobzi — happy to share our methodology with institutional data teams.
👉 GetJobzi.com — Talk to us