Skip to main content
Product

The Health-Tech Agency Scorecard: 12 Capabilities a Diagnostics Startup Actually Buys (With Our Own Scores)

A rubric a diagnostics founder can apply to any shortlisted agency, including ours, with the evidence to demand for each capability and the four we score low on and refer out.

2026-09-11 · By Filip Lauc

Why a diagnostics startup needs a capability scorecard, not a portfolio review

Because portfolios show what an agency shipped, not whether it has handled the specific mechanics a diagnostics product fails on: instrument result ingestion, consent versioning, clinician gating, kit logistics, and cross-system deletion. A scorecard forces each vendor to answer the same 12 questions with evidence rather than case-study prose.

A diagnostics platform is not a SaaS app with a health theme. It is a physical supply chain (kits out, samples in), a laboratory workflow (instruments, result files, re-runs), a regulated data flow (consent, controller/processor boundaries, deletion), and a commerce layer (subscriptions, B2B partners, multi-market shipping) that all have to agree on one identifier for one person's sample. Most agency portfolios contain none of that, and a beautiful marketing site plus a Stripe checkout tells you nothing about whether the team has ever had to explain a result-release rule to a clinical advisor.

The scoring scale below is deliberately blunt. 0 means never touched it. 1 to 2 means read about it or did it once under supervision. 3 means shipped it once and it held. 4 means shipped it repeatedly and it survived contact with auditors, partners, or an incident. 5 means we have operated it in production for years, broken it, and fixed it. We publish our own scores in the next two sections so the rubric is usable against us, not just by us.

The 12 capabilities, the evidence to demand, and Jaspero's own score

The 12 capabilities a diagnostics startup actually buys are lab and instrument result ingestion, sample chain-of-custody, consent capture and versioning, clinician-gated result release, kit batch and warehouse logistics, partner and B2B dashboards, controller-versus-processor DPA input, GDPR deletion across every system, payments and subscriptions for test kits, multi-market shipping, audit logging, and BI reporting.

For each one, do not accept a yes. Ask for one of four artifact types: a screenshot or screen share of the working feature in a real admin panel, a repo path or file walkthrough the engineers can narrate from memory, an incident record showing what broke and what changed after, or a named regulation clause the team can point to as the reason a specific design decision was made. If a vendor cannot produce at least one of those four for a capability, score it 2 at best regardless of what they claim.

Our scores below come from six years as GlycanAge's engineering team across 10+ interconnected systems, plus adjacent logistics and payments work on other products. Where we score ourselves a 3 rather than a 4 or 5, it is because we did it in one context and would not claim it generalises yet.

  • Lab/LIMS and instrument result ingestion. Evidence: a parser repo path, a sample result file, the re-run and outlier handling rule. Our score: 5. Built and maintained result ingestion from laboratory output into customer-facing reports for years, including format changes and re-analysis.
  • Sample chain-of-custody. Evidence: the event log for one barcode from kit dispatch to result publication. Our score: 5. Every state transition is an append-only event, which is what makes 'where is my sample' answerable by support rather than by an engineer.
  • Consent capture and versioning. Evidence: the schema showing consent text version pinned to a sample, not just a boolean on a user. Our score: 5. Consent text changes; results already issued under prior versions must remain attributable to the version the person actually signed.
  • Clinician-gated result release. Evidence: a screenshot of the review queue and the rule that decides which results require human sign-off before the customer sees them. Our score: 4. Shipped and operated, though the gating policy itself is always the client's clinical decision, never ours.
  • Kit batch and warehouse logistics. Evidence: batch and expiry tracking, pick/pack screens, stock reconciliation between warehouse and system. Our score: 4. Built warehouse and kit-batch systems in production; also transferable from our farm-to-table logistics work.
  • Partner and B2B dashboards. Evidence: the permission model for a partner who must see their customers' orders but not their raw results. Our score: 5. Partner dashboards with scoped access were a core part of the GlycanAge system set.
  • Controller-versus-processor DPA input. Evidence: the specific DPA clause that changed an architecture decision, named. Our score: 4. We have sat in these drafting conversations as the processor and changed data location and retention behaviour because of them. We are engineers, not your data protection lawyer.
  • GDPR deletion across every system. Evidence: a deletion runbook listing every system, including analytics, email tooling, backups, and BI copies. Our score: 4. Deletion across 10+ systems is the capability most teams discover is missing on the day the first request arrives.
  • Payments and subscriptions for test kits. Evidence: refund, partial-refund, and failed-sample flows, not just a checkout. Our score: 4. Kit commerce has an awkward edge: the money is taken before the deliverable physically exists.
  • Multi-market shipping. Evidence: carrier integrations, customs and documentation handling, and what happens when a sample is stuck in transit past viability. Our score: 4.
  • Audit logging. Evidence: who viewed which result, when, from where, and how long that log is retained. Our score: 4. Read access to results is the log auditors ask for and the one most products never implemented.
  • BI and operational reporting. Evidence: a live dashboard, and the answer to where the reporting copy of personal data lives. Our score: 4. We move reporting workloads to BigQuery or Postgres when the operational database stops being the right tool.

The four capabilities we score low on and refer out

We score 0 to 2 on CE marking and MDR software classification, FDA 510(k) submission support, on-premise hospital HL7 and FHIR interfacing, and clinical trial EDC systems. These are separate disciplines with their own auditors and specialist firms, and hiring a product engineering team to cover them is how founders end up with an expensive rewrite.

CE marking and MDR classification (score 1): we can build to a quality process someone else defines and we can keep technical documentation inputs current, but classifying your software as a medical device, running the conformity route, and dealing with a notified body is regulatory consultancy work. Ask us to implement the controls, not to decide whether you are Class I or IIa. FDA 510(k) (score 0): we have not supported a submission and would not pretend to. If your route to market is US clearance, your first hire in this area is a regulatory affairs consultant, and the software team follows their design controls.

On-premise hospital HL7 v2 and FHIR interfacing (score 2): we have built API integrations across many systems and we read the specs, but interface engines, VPN tunnels into hospital networks, and multi-month IT procurement cycles with individual hospital trusts are a distinct competence. Clinical trial EDC (score 1): eCRF design, 21 CFR Part 11 validated systems, and monitor workflows belong with an EDC vendor or CRO. When these come up we say so early and point clients elsewhere, because the alternative is billing for a learning curve.

The three agency profiles you will meet, and how each one fails

A diagnostics founder shortlisting vendors will meet three profiles: the generalist web shop, the MedTech or regulatory consultancy, and the embedded product team. Each fails in a predictable way. The web shop underestimates the physical and regulated layers, the consultancy produces documents nobody can ship from, and the embedded team is too expensive and too involved for a pure marketing-site scope.

The failure mode matters more than the price, because the cost of picking the wrong profile is not the invoice, it is the rebuild. A generalist shop will happily build your kit shop and customer portal, and it will look good. The failure surfaces at month nine, when a partner needs scoped access, a sample needs re-running, a deletion request arrives, and there is no event log to reconstruct anything from. The consultancy failure is the inverse: excellent documentation of what should exist, delivered to a team that still has to find engineers to build it. Embedded teams fail by becoming the only people who understand the system, which is a real risk you mitigate with code ownership terms and documented handover, not with hope.

  • Generalist web shop. Cheapest per hour, fastest to a pretty front end. Failure mode: no model for samples, consent versions, or cross-system deletion, so the data layer gets rebuilt once volume or an audit arrives. Fine for the marketing site, risky for the platform.
  • MedTech or regulatory consultancy. Highest day rate. Strong on classification, quality systems, and submissions. Failure mode: deliverables are documents and processes, not running software, and implementation gets subcontracted to whoever is available.
  • Embedded product team. Mid to upper rate range, retained monthly rather than fixed-bid. Owns the interconnected systems, on-call, and the boring maintenance years. Failure mode: key-person dependency and scope creep into decisions the founder should own. Overkill if you need one site and one checkout.

How to run the scorecard in a real vendor call

Send the 12 capabilities to each shortlisted vendor before the call, ask them to self-score 0 to 5 with one piece of evidence per line, and then spend the call verifying three scores at random by asking the engineers, not the salesperson, to walk you through the artifact. Confident self-scored 3s are worth more than uniform 5s.

Two questions separate teams that have done this from teams that have read about it. First: 'Tell me about a time a result was released that should not have been, or nearly was, and what you changed.' A team with real operating history has an answer, usually an unflattering one. Second: 'Walk me through your deletion runbook and name every system on it.' Counting systems out loud is hard to fake, and if the list stops at the database while analytics, email tooling, and BI copies go unmentioned, you have found the gap.

Then weight the scorecard to your actual next 12 months. A pre-launch diagnostics startup needs capabilities 1 through 5 and 8 at a high score and can defer partner dashboards. A company already selling through clinics needs 6, 9, and 10. A company entering the US needs a regulatory consultant before it needs another engineering vendor. We publish our own scores so a founder can put us on the same sheet as everyone else, decide where we genuinely fit, and route the four capabilities we are weak on to people who do them properly.

Key Takeaways

  • A diagnostics platform buys 12 specific capabilities, from instrument result ingestion and consent versioning to cross-system GDPR deletion and audit logging. Score every vendor on the same 12.
  • Accept only four evidence types per capability: a screen share of the working feature, a repo path the engineers can narrate, an incident record, or a named regulation clause that changed a design decision.
  • Jaspero scores 4 to 5 on the eight operational and data capabilities, and 0 to 2 on CE/MDR classification, FDA 510(k) support, on-prem hospital HL7/FHIR, and clinical trial EDC, which we refer out.
  • The three vendor profiles fail predictably: web shops miss the sample and consent layers, consultancies deliver documents instead of software, embedded teams create key-person dependency.
  • Weight the scorecard to your next 12 months. Pre-launch needs lab ingestion, chain-of-custody, consent, and gating. Clinic-channel growth needs partner dashboards, payments, and shipping.

The scores above come from six years as an embedded engineering team, documented in our GlycanAge case study, where we built and maintained 10+ interconnected systems across lab, warehouse, customer, and partner surfaces.

Frequently Asked Questions

What should I ask a software agency before hiring them for a diagnostics startup?

Ask them to self-score 0 to 5 on the 12 capabilities a diagnostics platform actually needs, then verify three scores at random with the engineers rather than the salesperson. The two highest-signal questions are 'tell me about a result that was released or nearly released when it should not have been' and 'name every system on your GDPR deletion runbook'. Teams with real operating history answer both immediately.

Do I need a regulatory consultant and a software agency, or can one vendor do both?

In most cases you need both. Classification under MDR, notified body interaction, and FDA submissions are regulatory disciplines with their own specialists, while product engineering implements the controls that process defines. An agency claiming to cover both should be asked for a named submission or conformity route it has actually supported.

How much does it cost to build a diagnostics platform with an agency?

It depends far more on how many of the 12 capabilities are in scope than on hourly rate. A marketing site plus a kit checkout is a small fixed-bid project; lab result ingestion, chain-of-custody, consent versioning, clinician gating, warehouse logistics, and partner dashboards form a multi-system platform that is normally delivered as a retained team over quarters, not weeks. Scope the capability list first, then ask for pricing against it.

What is the difference between a controller and a processor when hiring a health-tech agency?

The startup is almost always the data controller, deciding why and how personal and sample data is processed, and the agency operating or building systems on its behalf is a processor. That split drives your DPA, sub-processor list, data location, breach notification timelines, and deletion obligations. Ask a prospective agency to name one DPA clause that changed an architecture decision on a past engagement.

Can a generalist web agency build a diagnostics product if we handle compliance ourselves?

For the marketing site and a simple kit shop, yes. For the platform, the risk is that compliance is not a document you apply afterwards but a data model: consent pinned to versions, samples modelled as event streams, audit logs on result reads, and deletion that reaches analytics and BI copies. Retrofitting those into a system built without them usually means rebuilding the data layer.

Filip Lauc

Written by

Filip Lauc

CEO, Jaspero

Filip Lauc is the CEO of Jaspero, a software development agency based in Osijek, Croatia. A full-stack JavaScript developer with over a decade of experience across Angular, Svelte, and Node.js, he leads Jaspero's work as a long-term embedded engineering partner for clients like GlycanAge, where his team has served as the dedicated engineering team for six years.

Let's Build Together

Your vision,
our expertise.

From AI integration to full-stack development, we turn ambitious ideas into products that perform.