Enjins · Part 1 of 2
Ten Years of AI Engineering, Ten Years of AI Tech Investments
Nick Jetten and Bastiaan van de Rakt on a decade of production ML & AI and AI tech investing — VC/PE convergence, AI-native funds, TechTruth benchmarks, and five findings from 100+ tech due diligences.
This article was originally published on Enjins. The content, photographs and links below follow the original publication.
A two-part whitepaper by Nick Jetten and Bastiaan van de Rakt
Nick Jetten and Bastiaan van de Rakt are co-founders of Enjins and Deeploy; Nick is CEO of Enjins. Bastiaan invests through Why Commit Capital, in collaboration with Volve Capital and AENU, among others. Both write here in a personal capacity, sharing learnings from the last ten years.
In one line: for ten years we have run two jobs in parallel, building ML & AI systems that had to survive production, and backing the companies that build them. The first job is what made the second one work. Every judgement in this paper, and the whole diligence instrument in Section 3, came primarily out of engineering practice rather than investment theory. That is the argument: as AI engineering moved from can we build the model to can we operate the system, only people who run AI companies can separate the real ones from the convincing fakes.
Why we are writing this
We met twelve years ago in the beginning of the predecessor of the AI (then called machine learning) practice at VODW (now part of EY) building some of the first genuinely scalable ML applications for large enterprises. It did not take long to work out that the interesting version of the problem was somewhere else: within a few years our full focus had moved to startups and scale-ups, which is where the real AI revolution actually started, and where it has stayed.
That gives us an unusual vantage point. Enjins has helped more than 100 organisations move AI from prototype to production since 2018. Deeploy, which we co-founded in 2020, exists because “the model works” and “the model is accountable in the real world” are two completely different sentences. Together, we have conducted over 100 AI audits and made 25 direct investments. Across the decade we have seen roughly 7,000 decks, with an increase of “AI-first” or “AI-driven” claims year by year; about 1,500 of them have been run through the current TechTruth pitch deck checker methodology end to end. The double human feedback loop on top of the pitch deck checker is what makes the feedback on every new check more relevant and accurate. Double, because we check the outcome against both our own judgment as well as the investors and/or founder judgment.
Part 1 – What Actually Changed
Over the last ten years, we have experienced first-hand how AI is fundamentally rewriting the investment landscape. In Part 1, we unpack this shift across three key themes:
- The Changing VC/PE Landscape: How the historical boundary between venture capital and private equity has effectively dissolved.
- The AI-Native Investor: why and how funds are rebuilding their own operations to become AI-native, starting at the front of the funnel.
- Executing Tech DDs in the AI Era: How AI impacted technical due diligence, both in WHAT you need to assess as well in HOW you need to assess this.
As a bonus we will share key lessons learned from our tech DD work.
In part 2 we look beyond the wire transfer. We will focus on value creation post-deal and look ten years ahead instead of back. Discussing the future of AI & Tech investments and AI-native investing 2.0.
Section 1 – The investment landscape is being rewritten
1.1 The category boundary between VC and PE has effectively dissolved
For thirty years the division of labour was clean. Venture capital bought optionality on things that did not exist yet. Private equity bought cash flows that did exist and improved them with leverage and operational discipline. The two rarely competed for the same asset.
That boundary has effectively collapsed, wiped out by the shift toward AI-integrated operations.
Let’s look at a tangible example: General Catalyst, a giant tech venture firm with $30bn under management, proves that line has disappeared. They raised $1.5bn for a strategy that built an acquisition company called Long Lake, which has bought around 30 businesses since 2023. Most notably, they backed a $6.3bn deal to take American Express Global Business Travel private; a massive, profitable corporation making $2.7bn in revenue and $182m in operating profit. That is a classic Private Equity buyout, executed by a Venture Capital firm.
Thrive, 8VC, Khosla, and Slow Ventures have all moved in comparable directions in the U.S. Meanwhile, Europe is beginning to follow. Berlin-based players like Aven Capital or Tenet Capital were specifically launched to execute “AI roll-ups”, a strategy of acquiring traditional, cash-flowing service businesses and overhauling their operational margins using a shared AI software platform. Sweden’s EQT, historically a heavyweight Private Equity player, launched EQT Ventures for early-stage VC bets while deploying its proprietary Motherbrain AI engine across its entire portfolio to bridge the gap between venture tech and buyouts.
Instead of asking “Is this a VC deal or a PE deal?”, the question that actually matters today is: “Who is actually going to transform how this business operates with AI, and do they have the engineering track record to pull it off?”
1.2 Roll-Up 2.0: valid thesis, wrong timeline
The classic buy-and-build (or roll-up) strategy always relied on the same two levers: multiple arbitrage and a centralized back office. Buy ten regional occupational health providers or mid-sized precision machining facilities at 4x, sell the platform at 10x, centralise finance and procurement along the way. There was real added value in doing so, but the upside was bounded, and it did not require the acquirer to be technically capable.
Roll-Up 2.0 makes a much stronger claim: that AI can compress the delivery cost of a services business (i.e. legal review, claims handling, bookkeeping, call centres, field scheduling) by a factor large enough to change the margin structure of the whole category, not just the overhead line. When AI legal platform Eudia acquired alternative legal service provider Johnson Hana, they pursued a strategy far deeper than simply combining law firms to trim administrative overhead. AI handles the heavy lifting, such as initial intake, clause extraction, and risk tagging, while a drastically smaller team of oversight lawyers approves the output, turning a traditionally low-margin legal workforce into a high-margin, software-driven factory.
These examples are appearing rapidly: Crescendo in contact centres, Percepta in asset management. These are all bets that the gross margin of a labour business can be dragged toward the gross margin of a software business.
We agree that this thesis is directionally correct. We also think the execution risk is systematically under-priced; investors are paying high valuations and making big bets without factoring in how often these projects fail or run over budget, for a reason that is obvious to anyone who has actually shipped these systems:
The bottleneck is never the model. It is the data estate, the process variance, the system variance, the needed change management and the accountability layer.
A ten-site services business does not have one process. It has ten processes with a shared name, each encoded in a different practitioner’s head, sitting on four incompatible legacy systems of record, with no ground-truth labels anywhere.. Building the model takes two weeks. Getting that model to consume reliable data, integrating it into the daily workflow, and winning the trust of the experts who use it? That is an eighteen-month project. Underwriting the former and paying for the latter is the single most common structural error we see in Roll-Up 2.0 models.
A second, subtler error: the acquired margin is often not defensible. If your edge is “we applied a frontier model to a services workflow”, your competitor can apply the same frontier model next quarter for the same price. The defensibility has to come from the proprietary data exhaust, the domain-specific feedback baked into the setup, the switching cost of the workflow, or the regulatory/audit position, not from the intelligence itself, which is a rented commodity with a falling price.
1.3 AI operators: the new scarce input
Ten years ago the scarce input in a fund was proprietary deal flow. Five years ago it was brand. Today, in the segment of the market we live in, it is operators who can tell the difference between a demo and a system, and who know what is needed to build such a system.
This is the quiet reason venture firms are hiring engineers, and PE firms are building in-house AI functions rather than buying advisory hours. The knowledge is not transferable through a deck. It is transferable through someone who has actually spent years in the trenches building software and shipping AI projects into production.
The investor world has priced this. It has not yet staffed it, which is precisely the gap we are building into.
1.4 The AI-native fund is not a fund that uses AI tools
Almost every fund now has Affinity or a similar CRM, a note-taking agent, and one or two appointed “AI Champions” or working students launching initial workflows. Yet, even where ambition is high, especially in larger teams, funds consistently struggle to get everyone aligned around a single AI-native way of working.
The typical maturity state we see is a fragmented one: ad-hoc personal usage of desktop apps, maybe one or two shared Claude skills used by half the team, one custom web tool deployed, and pervasive doubts about what is actually compliant. Shadow AI spreads rapidly (good that there’s software for that). Then, right before a team-wide workflow gains traction, a major AI model release drops, making partners skeptical about investing further time: “Should we just wait a little bit longer?”
No, you should not. You need to design an overarching, centralized way of working with AI. One that compounds lessons over time, stays sovereign and compliant by design, and is supported by serious investment so that adoption becomes frictionless. Transitioning a fund of twenty investment professionals into true AI-native investors requires dedicated capital and operational commitment, not the hope that an intern or working student can carry the transformation alone. In the next section we explain how it can be done.
Section 2 – How PE and VC funds actually become AI-native: the front of the funnel
A fund is, structurally, an information-processing organism with four organs: sourcing, diligence, portfolio value creation, and LP relations (fundraising and reporting). Becoming AI-native means rebuilding each organ so that your scarce resource (your investment professionals) can dedicate their time exclusively to human judgment, while everything upstream (the repetitive, manual tasks of gathering, organizing, and visualizing data) is entirely machine-handled.
We will walk through the four organs, and give two key pre-requisites in becoming AI native.
2.1 Sourcing — from coverage to prioritisation
The wrong goal is to see more companies. The right goal is to reject faster, on stated grounds, before anyone has spent attention on the company.
Every fund is at risk of drowning, and most are already there. Our own intake is the mirror: TechTruth has scanned 1,021 decks across 12 industries (date 4 July 2026), 901 of which calibrate the benchmark, with 374 human verdicts fed back into the scorer. Two findings recur at that scale. 61% of decks describe an AI product with no proprietary AI asset underneath it. 49% contain at least one founder claim that does not reconcile with public sources. Both are machine-checkable, and both are checkable before a first meeting.
A high-performing sourcing function starts with a written, versioned thesis. From there, you build an automated first-pass layer to apply that thesis across the funnel and crucially, you close the loop by feeding human verdicts back into the scorer to sharpen its accuracy.
The first two are common. The third is what almost nobody builds — funds included, which is awkward, because it is exactly the thing we mark companies down for in Finding 5. A flywheel on a roadmap is a claim; 374 recorded TechTruth verdicts are a real asset. That is why they are worth more to us than the next thousand decks.
2.2 Diligence — machines verify, humans judge
The single highest leverage change available to a fund today is to stop using humans for verification and start using them exclusively for judgement.
Verification is: does this founder’s claimed role at their previous employer match the public record? Is the architecture diagram consistent with the job postings? Does the commit history support the claimed engineering velocity? Are the unit economics arithmetically coherent at the stated token cost? These are mechanical, high-volume, boring, and machines are now better at them than tired associates at 11pm.
Judgement is: given that this team has thin technical depth but an extraordinary distribution wedge, do we believe they can hire in? Machines are not close, and pretending otherwise is how funds get hurt.
The productivity gain is not “diligence gets faster.” It is that the verification floor rises to 100%. Today, in most funds, deep technical verification is applied to maybe the final 5% of the funnel — which means the other 95% is rejected or advanced on narrative quality alone. That is a systematic bias toward good storytellers, and our data says the bias is enormous.
Across our benchmark, founder/team scored 6.9/10 while technical moat scored 5.5, infrastructure 4.8, and AI asset depth 4.7. The narrative is consistently running two full points ahead of the substance.
That two-point spread is not a description of the market. It is measuring how much investors are willing to overvalue the story over tangible technical execution.
Those are the two organs that operate before any money leaves. The two below only start working once you own the asset, and they are where most funds have the furthest to go. We treat them properly in Part 2; here is the shape of them.
2.3 Portfolio value creation — the honest version
Most funds’ “AI value creation” offering is a slide deck and an introduction to a vendor. The honest version has three levels, and the level matters more than the ambition.
Adoption: the portfolio company’s people use AI tools well. Cheap, fast, real but small. Typically stays on improving individual efficiency levels. Weeks.
Workflow: economically significant processes are rebuilt around AI with measured before/after unit economics. Months for single workflows. Quarters to automate a serious % of the total manual work in large ops or sales teams.. This is where the actual EBITDA lives.
Product: the company’s own offering becomes AI-native, changing what it can sell. Quarters to years, and only worth attempting where the data estate supports it.
The failure mode is funds promising level 3 while their portfolio companies are not ready for level 2.
2.4 The internal knowledge estate
The unglamorous one, and possibly the highest ROI. A fund accumulates an extraordinary proprietary corpus: every memo, every rejected deck, every board pack, every post-mortem, every opinion in the Monday morning internal “top-deals” call. In almost every fund we have seen, this corpus is functionally write-only. Nobody can ask it a question. And all the information (and judgment!) that went through the sourcing and dd process is only partially stored.
Making that corpus queryable — with proper access control, provenance, and the ability to answer “what did we say about this category eighteen months ago and were we right?” — turns institutional memory into an asset. It is also, not incidentally, the precondition for any of the sourcing and diligence automation above to be calibrated on your own judgment rather than a generic model’s. Typically, we see that the operational systems that should jointly spin up this corpus already exist; a modern CRM, cloud file storage, and a few key external sources. But it’s messy, disconnected, and unstructured. Throwing a data lake, vector database, or new knowledge base on top of a mess solves nothing. Start by structuring and tagging your data so that a new human colleague could actually find and understand it. If a human can’t make sense of it, an AI agent won’t either.
2.5 The most important prerequisites; centralize, invest, govern.
Centralizing your data corpus is only step one; you must also centralize and streamline how your team accesses AI tools. You can achieve this by investing in a custom, centralized interface (e.g., yourfund.com/AI) accessible via browser, with all necessary data connectors pre-installed. Crucially, this requires building a serious, compounding AI harness over time: a software foundation that continually captures learning, integrates new workflows, and grows in value, which ultimately demands dedicated internal development capacity or an external engineering partner. This setup allows investment professionals to be users of a finished system rather than expecting them to act like semi-developers.
Note that “skill” undersells what actually has to be built. A skill is a prompt-level instruction, and it behaves differently for every person who invokes it. The agents we build are tailored workflows that combine deterministic code with LLM calls, so the reliable steps stay reliable and the model is only used where it is genuinely needed. Same input, same output, every associate, every time.
Where you run it matters just as much. Every fund we speak to is nervous about exposing its own corpus, the memos, the verdicts, the board packs, to a third-party model. Hosting the agents in your own private cloud resolves that: the proprietary knowledge becomes available to the workflow without ever leaving your perimeter.
No matter your technical setup, governance must be built in from the start. This is the entire reason that Deeploy exists. If you skip this, you guarantee one of two outcomes: either your Chief Compliance Officer eventually shuts everything down and sends you back to square one, or you force your team into shadow AI usage. Because yes, associates will use unvetted desktop models to check deal data behind your back.
Section 3 – Executing Tech DDs in the AI era
We started exactly where everyone starts: a two-hour call with the CTO, an architecture diagram on a screen share, and a gut feeling written up as a memo. It was not bad work, but it was unrepeatable. Unrepeatable work cannot be compared, and work that cannot be compared cannot be benchmarked. The methodology below is what ten years of that frustration turned into.
Roughly 2016–2019: from vibe to checklist. The first real improvement was simple ; writing down the questions in advance and asking every company the same ones. That became a structured CTO survey, now around 38 questions plus a standing document request. The immediate discovery was that we had been asking sharp questions of the companies we already doubted and soft questions of the ones we liked.
Roughly 2019–2022: from checklist to modules and scores. Consistent questions produced comparable answers, which made scoring possible. We settled on four modules — Product & Tech Roadmap, Infrastructure, Tech & AI Moat, Tech & AI Team — each with five dimensions, each scored 1–5, where 1 is severely underdeveloped and 5 is best in class.
Roughly 2022–2024: from scores to a benchmark. A score in isolation is a number with an opinion inside it. A score against 100+ comparable diligences is a finding. Once the corpus was large enough, we could tell a board not merely that their test coverage was low, but which of their weaknesses were normal for the stage and which were genuinely anomalous.
Roughly 2024–2026: from benchmark to automation, because AI outran the manual process. Walking through a pitchdeck was a manual exercise; checking a code base needed multiple hours of an engineer, and then still it was hard to see everything.
So we automated. On the deck side that became TechTruth, with human verdicts fed back to recalibrate the scorer. On the diligence side, the last 15 tech & AI DDs have been roughly 90% automated and fully benchmarked — including an automated code review that reads the repository directly and reports code structure and quality, architecture and maintainability, security and robustness, testing and reliability, AI readiness, and IP defensibility.
Now, in the interviews, we can focus on the flags found in the automated analysis instead of starting from scratch. This is also the part we will not automate, and we do not want to. It’s the same as the investor’s judgment that a fund should not automate; the judgement call we make in DDs on whether a team can grow into the gap between what they have built and what they have promised is something that should not be the answer of a prompt.
One design choice we would defend above all the others: a flag taxonomy with time attached, not just severity. Red = a direct blocker. Orange = attention within three months. Yellow = attention within twelve months. Green = a genuine strength, recorded with the same rigour as the risks.
Bonus: five findings across our tech & AI DDs
This is the section we argued about most, because several of these findings cost us money to learn.
Finding 1 — The pitchdeck says proprietary AI agent. The code says wrapper.
Claim: the vocabulary of AI differentiation now moves roughly four times faster than the architecture underneath it.
61% of scanned decks show wrapper characteristics with no proprietary asset; 32.7% are pure wrappers — a workflow layer with no defensible architectural differentiation. Between 2025 and 2026 “AI-powered” became “agentic” across the market without a corresponding change in what was being built.
The difference between a wrapper and a real agentic system? A real agentic system has a recovery path when a tool call times out, a mix of deterministic code and LLM calls, an audit trail and in the best cases a mature LLMOps architecture that allows fast and automated iterations. A wrapper has an ordered list of LLM calls and a router. The value for the user might look identical; the failure behaviour does not.
The point is not that wrappers are uninvestable — we score the honest version of one as a perfectly fundable 3 out of 5. The problem is never the architecture. It is the gap between the architecture and the promise in the pitchdeck.
The test: run a technical code review to see whether a gap exists between the pitch deck and the actual codebase. If you find one, call it out directly to the founders. If they candidly admit the limitation and explain their roadmap toward true durability, you have an honest team you can price accurately.
Finding 2 — The cheapest check in venture is the one nobody runs
Claim: 49% of decks contain at least one founder claim that does not reconcile with public evidence — and almost every fund discovers this after the third meeting rather than before the first.
Mostly this is not fraud. It is rounding, title inflation, a project described as owned that was contributed to, a headcount that includes contractors. Individually trivial. In aggregate it is the single most predictive cheap signal we have, because the pattern of small inflations tells you how the team will describe reality to a board.
The lesson is about sequencing, not suspicion. Verification is mechanical and costs minutes, so it belongs at the top of the funnel — where it changes which meetings you take — rather than at the bottom, where it only changes a decision you have already emotionally made.
The test: automatically check a founder’s résumé and claims against public records before your very first phone call, not right before signing a term sheet.
Finding 3 — Charisma scores as capability, and it is measurable
Claim: founder quality is genuinely predictive, which is exactly why it contaminates everything else.
A strong founder impression leaks into the technical assessment and quietly raises every other score. Every scoring framework we have built has had to structurally suppress that leak to stay useful: Tech & AI Team is scored as its own module, with its own dimensions, in which we interview multiple team mates and not just rely on the founder or the CTO. A strong team moreover tells you that the founding team is capable of hiring the talent they need.
One dimension in that module has become disproportionately informative: how the engineering team uses AI on itself. A team that has adopted coding agents deliberately — tried autonomous merge-request agents, rejected them for stated reasons, written down guidelines, embedded it across the whole tech team — tells you more about its engineering judgement than any answer to a question about strategy.
The test: go beyond the founding team in the DD, which ensures a high quality tech standard and ensures that the founders are capable of finding (and retaining) a top tech team.
Finding 4 — The Ghost Scan: one author, a lot of generated code, and nobody who can maintain it
Claim: across the last twenty-odd Ghost Scans Enjins has run, the recurring finding is not bad code. It is a codebase far larger than the team that has to live with it, with effectively one author.
The Ghost Scan reads the commit history: who wrote what, in what volume, and how much of it a second person has ever touched. The pattern repeats ; one engineer is the only meaningful contributor to the core paths, output per engineer-week sits far above what a human writes unaided, and large blocks are visibly model-generated.
Generated code is not the problem; we would be more worried by a 2026 team not using coding agents. Volume is. A model writes a plausible implementation for every case you describe, so the estate grows to cover cases instead of being designed down to abstractions. It runs, it demos, and no second human has ever reasoned about it end to end.
Two consequences. It is not robust; a system nobody designed fails in ways nobody anticipated, usually in production. And it is a key-person risk: one person generated the estate and nobody else has read it, which is not an engineer to retain but a single point of failure with a notice period. That does not make a company uninvestable. It makes the timeline wrong: the team needs one deliberate consolidation quarter before it can absorb hires at the promised rate.
The test: pull the contributor graph for the core modules and ask what share has been touched by exactly one person. Then scan the codebase for hard-coded credentials, endpoints and thresholds, and ask the CTO which were meant to be configurable. The gap is the size of the consolidation quarter.
Finding 5 — Most moats are raw material, and the ones that hold are increasingly physical
Claim: the data flywheel is usually on the roadmap but actually non-existent — and the easier it becomes to write software, the less software by itself can defend.
The pitch is never short on flywheel language. What the diligence almost never finds is the pipeline that would make it true. Where does the feedback actually live, and in what form? Is there a real mechanism — not a Slack channel someone means to get to — that routes a customer correction back into a retraining set, an eval case, or a prompt change? A flywheel plan is a claim. A flywheel with eighteen months of dated, checkable history behind it is an asset. So don’t invest or underwrite based on good intentions only; underwrite the conversion, with dated milestones and a named owner, not the intention.
Meanwhile, the highest defensibility scores we award rarely go to the most complex agent architectures. They go to teams that control a structural mechanism to generate data nobody else can access; for instance through owned physical infrastructure or deeply embedded, exclusive partnerships with corporates. In recent DDs in climate and energy we saw this clearly with proprietary data flywheels: a battery operator generating minute-by-minute grid telemetry from its own fleet, or a software platform with exclusive data-sharing rights embedded deep inside a utility’s operations. That raw operational data cannot be bought at any price because it does not exist on the open market. Stack a robust software layer and a well-architected AI engine on top of an exclusive data pipeline, and multiple defensible layers compound rather than just one.
The test: ask to see the data flywheel in a database, what is the timestamp on the first record and the last one? And ask what the company owns that a well-funded competitor could not replicate simply by writing more code.
What the last ten years actually taught us
Strip it back and there is one sentence underneath all of this.
In 2016 the scarce thing was the model. In 2026 the model is a commodity you rent by the token, and the scarce thing is proof — proof that a system works outside the demo, proof that the data is really yours, proof that AI adoption drives measurable EBITDA expansion, proof that someone is accountable when it fails.
Capital is repositioning around that scarcity. All three of the shifts in this paper are the same shift, seen from different distances.
Part 2 — coming next
Part 1 stopped at the wire transfer. Part 2 starts the day after it clears, and it has the uncomfortable half of the material: the findings that never appear in a deck and always appear in the integration plan — why infrastructure debt is always the same five items, why AI operations are still the binding constraint on iteration speed at scale, and why a feedback loop is not a claim until it runs itself.
And then we look forward. Part 2 takes the same lenses ten years out: what happens to Roll-Up 2.0 when the AI margin advantage is fully competed away, what happens to a fund when diligence stops being an event and becomes a continuous function, and what the job of an investor actually becomes once verification is free. Where Part 1 is evidence-led, Part 2 is prediction-led — so every major claim will carry a date and a stated way of being proven wrong.
A note on what we are building
Everything above is drawn from work we do: Enjins builds production AI systems and runs tech and AI due diligence for VC and PE; WhyCommit — in collaboration with Volve Capital and AENU — invests at pre-seed and Series A in AI and climate; TechTruth.ai is the deck-verification layer that produced most of the benchmark data in Sections 3 and 4; Deeploy handles the accountability layer that Part 2 argues you cannot skip.
We are also building out a dedicated PE/VC practice; Enjins Investor Services. With people who can sit on both sides of the table, running technical diligence one week and rebuilding a portfolio company’s data estate the next. If you have shipped AI systems into production and you find capital allocation interesting, that combination is rarer than it should be and we would like to talk to you.
Reach us via enjins.com or directly through Nick and Bastiaan on LinkedIn.
Acknowledgements
This paper is written in our two names, but almost none of what is in it was learned alone. Without our co-founders Maarten Stolk and Tim Kleinloog we would not be anywhere near this point in the journey — the instrument in Section 3 exists because they built the companies around it with us, and argued with us about it along the way.
And our thanks to Joanne Lijbers, Anouk Meerman, Yannick Hermsen and Bart Maassen, for input, for support, and for writing along with us. Several of the sharper formulations in this paper are theirs.
All benchmark figures in this paper are drawn from the TechTruth.ai deck corpus and work from Tech DD and Investor Services of Enjins as published, and reflect a beta-stage indicative scoring system, not formal due diligence. External market figures are sourced inline.