Observability Engineering second edition out now! 27 net-new chapters written for today's observability challenges.Get your copy

From 93% to 99%: hipages Turns a Decade of Blind Spots Into a Record Week

Australia's largest home services marketplace spent years trying to solve two issues: why job postings failed, and why tradies missed leads. Honeycomb’s distributed tracing gave them answers in weeks.

93% → 99%

Job posting success rate

10 min → 1 min

Spread in tradie notification delivery

3-6 mo → 1-2 mo

Frontend migration timeline

hipages logo
hipages icon
About

hipages is Australia's leading marketplace connecting households to home service businesses, known locally as tradies.

Industry

Online marketplaces / Home services

Use Cases

Distributed tracing, Frontend Observability, SLOs, Honeycomb Intelligence

hipages has been running its marketplace for 22 years, which means its engineering team maintains code written across several eras of the internet, in PHP, TypeScript, and Go, some of it older than some of the engineers now maintaining it. Two separate problems had been hiding under one general sense that something wasn't right. The first: a straightforward defect whose true impact nobody had grasped, and the second: a decade-old complaint nobody could pin down. What changed was a decision to instrument the stack end to end with OpenTelemetry and put Honeycomb behind it.

Before Honeycomb

hipages had the tooling most companies accumulate over two decades: plenty of logs, plenty of metrics, none of it connected. Engineers could tell at a macro level when something was broken and could often tell, service by service, whether a given piece was technically healthy. What none of it could show was the customer's actual experience moving through the system, because there was no way to follow one request across the dozen services it touched on its way from a homeowner's browser to a tradie's phone. The topline numbers looked fine because topline numbers hide the kind of problems hipages actually had.

Jeremy Burton, hipages' CTO, needed something that could turn "we think this is a problem" into a number. The team rolled out OpenTelemetry instrumentation across the marketplace and trading platform and put Honeycomb behind it, starting with the basics: making sure trace and event data reached the tool before trying to do anything clever with it.

Problem one: the job posting form

hipages knew the job posting form had problems, but not how bad it was in aggregate, because the failures showed up as scattered, low-volume error reports, indistinguishable in a dashboard from ordinary noise. And each individual issue looked too small on its own to justify the cost of a dedicated investigation: chasing down one obscure failure mode for a handful of users was never going to be worth an engineer's time, so the reports piled up unaddressed rather than unnoticed.

Two weeks into instrumenting, the SLO on the job posting flow surfaced a 93% success rate on the form that drives hipages' revenue-generating activity, the single most important interaction on the platform. Seeing the number was the first shock: nobody had realized 7% of submissions were failing outright. The second was seeing why: there was no single outage or smoking gun, but many small, distinct error conditions stacked on top of each other, each responsible for only a sliver of the missing jobs. That's what had made it too expensive to investigate before: every individual cause really was minor in isolation, and it took seeing them all ranked side by side to realize how much they added up.

This is where Honeycomb's BubbleUp did a lot of the work. Instead of engineers guessing at causes or grepping through a mountain of logs, BubbleUp let the team slice the failing submissions against everything else and automatically surface which dimensions correlated with failure: specific error conditions, steps in the flow, request shapes. That turned their 7% loss into a ranked list of concrete causes, so the team could go after the ones responsible for the biggest share of the loss first.

Working down that list over a matter of weeks, the team brought the success rate from 93% to over 99%, a lift that flowed straight into job posting volume, since more successful submissions meant more leads for tradies to claim. Joe, a principal software engineer who has been at the company for more than a decade, noted that the week the fixes shipped was the highest job posting week the company had ever recorded. It was the kind of increase in revenue that would have taken the growth team months to achieve.

Problem two: notifications arrived late

For more than a decade, some tradies had reported that by the time a job lead reached them, it had already been claimed. In hipages's marketplace, only the first few tradies to claim a lead get the contact details, so for jobs with high demand, a notification that arrives even a few minutes late might as well never have arrived. hipages had investigated the complaint multiple times, sometimes spending weeks trying to trace the path from a homeowner's job post through matching, scoring, and pricing, all the way to a push alert landing on a tradie's phone.

Every one of those investigations ran into the same dead end, and reached the same reasonable-looking conclusion. The individual services in the pipeline reported healthy average and p99 latency, and messaging lag was acceptable. Since nothing in hipages' own stack looked slow, the standing assumption was that any delay had to be happening somewhere the team didn't control: third-party push notification services, mobile carriers, or a tradie's own device and battery settings. That assumption was never really tested, because no tool could follow a single request through to the end, including the leg that disappears into the mobile network and the device itself, and nobody was measuring the number that mattered given how the marketplace works: the spread between the first tradie notified about a lead and the last.

With Honeycomb in place, the team instrumented that full pipeline (REST endpoints, database outboxes, Kafka) and extended trace propagation all the way through push notification dispatch to the moment it arrived on a tradie's device. That last leg, tracing a request from a backend service onto a physical phone, is the part Joe still calls the coolest thing the team has built with the tool. It converted a decade-old black box into a question with an answer, which let the team look at the actual distribution of end-to-end delivery times instead of per-step averages.

The distribution told a different story than the per-step averages had. The gap between the first tradie notified about a lead and the last one notified added up to multiple minutes in the worst cases. Some of that was outside hipages's control, like mobile carriers and device battery settings. But a meaningful share of it turned out to be inside their own stack after all, in specific services that per-step averages had been hiding. Once the team could see exactly where, the fixes were mostly configuration changes and added capacity, and they brought delivery times down from multiple minutes to about one minute in the worst cases. A month after the fix shipped, the notification complaint dropped off the top of hipages' tradie NPS feedback, and the team put monitoring in place so a regression wouldn't go unnoticed again.

Running the new site next to the old one

hipages is an SEO-dependent business, so migrating the website's frontend to a new technology stack carried real risk. Previous migrations of this kind had dragged on for months, largely because the only way to validate a change was to ship a small piece, wait, and watch real-world search rankings and traffic for weeks before moving on to the next step. This had killed most technology rewrites before they were finished.

For this migration, the team instrumented real user monitoring on both the legacy frontend and the new one, and ran them side by side, comparing actual visitor telemetry rather than synthetic load tests. That gave engineers and non-technical stakeholders the same empirical picture at the same time: line-by-line comparisons of real performance and SEO-relevant metrics between the two systems. The migration that would have taken many months, if it succeeded at all, was done in two months.

Every engineer gets to be Joe

Ten years of troubleshooting turns you into an institution, and that's exactly what had happened to Joe. He was the person other engineers called when something broke, because he'd accumulated more mental context about the system than any document could hold. That's useful until it becomes a bottleneck.

Earlier this year, hipages adopted Honeycomb's MCP server, which lets people ask questions about system behavior in plain language through Claude instead of writing queries by hand. Active Honeycomb users more than tripled in a month, from roughly 25 to about 80, and the growth wasn't confined to engineering. Product managers and support staff started using it directly, asking things like how fast a particular page was loading or what was happening for a specific customer, without opening a ticket with engineering first. The knowledge that used to live mostly in Joe's head now lives in traces anyone on the team can query.

The clearest evidence of that broader change came from a bug Joe couldn't crack. For five years, he had personally hunted a memory leak on hipages' SEO-critical directory pages, a leak that had been "fixed" with a restart of the process every few days. Last year, an engineer who had joined the company far more recently used Honeycomb and Canvas to dig into the details. He found the leak and fixed it in just a few days.

From knowing how the system works to deciding how it should work

Engineers spend less time reverse-engineering how the system currently behaves and more time deciding how it should behave. SLOs and alerting now sit in front of most problems, so the team frequently fixes an issue and tells the business about it after the fact, rather than waiting for a support ticket to start the investigation.

Job posting volume and the revenue tied to tradie credit consumption both grew off the back of fixes that used to be invisible, and the team no longer needs to grind out incremental funnel improvements to hit the same growth targets. Product and support staff can now answer whether something is broken for a given customer themselves in the time it used to take to file a ticket and wait for an engineer.

Without a tool like Honeycomb, it’s kind of like walking at night with a torch. You can shine it here and see one bit, then shine it there and see another. Honeycomb lights the whole stage, and all of a sudden I can see everything playing together, rather than just each individual piece on its own.

Jeremy Burton

CTO

Advice from Jeremy Burton, CTO

  1. Start small and instrument something real before you try to instrument everything.
    The value shows up fast enough to justify the next step on its own.
  2. Budget for the unglamorous part twice.
    Once to get telemetry flowing, and again for the slower work of adding business context to it. The second round is where the actual return shows up.
  3. Distrust your averages.
    The bugs costing you the most money are usually the ones an aggregate dashboard hides.
  4. Give the whole team access, not just your most senior engineer.
    Institutional knowledge that lives in one person’s head is a risk you’re carrying whether you’ve noticed it yet or not.