Alexander J. Perrin/ Working Theory

Before AI Had a Chat Window

Fifteen years of working with machine intelligence—from model operations and training-data integrity to private inference, agentic development, and the stubborn problem of keeping a human being in the middle of it.

Original essay for this site

A structured field of machine signals flowing into a narrow orange inference point and a command window
Editorial artwork for Working Theory · Model operations before the chat interface

I have worked with artificial intelligence long enough to remember when saying so out loud made people suspicious of you.

Not impressed. Suspicious. In 2012, "we use AI" was roughly as credible a business claim as "we use synergy," and it was frequently deployed by the same people.

My experience did not begin with generative AI. It began around 2011, when I was helping build the display and programmatic practice at iCrossing—a search agency expanding into a channel then described simply as "display," which would shortly become the modern programmatic advertising industry.

The vendor market surrounding that expansion was a circus.

Every week brought another demand-side platform, data-management platform, audience company, optimization engine, measurement provider, or media technology, each with a deck explaining why the previous eleven were wrong. Attribution was inconsistent. Definitions were negotiable. Margins approaching 90 percent were not unheard of, and were rarely disclosed. The industry had discovered that sufficiently complicated technology is indistinguishable from a moat.

I evaluated roughly thirty technology partners during this period. Most of them I could take apart inside an hour.

One I could not: Rocket Fuel.

Rocket Fuel claimed that artificial intelligence could evaluate enormous numbers of signals, predict which advertising opportunities were most likely to produce a desired outcome, and improve continuously as new information entered the system.

I assumed they were cheating.

So I went looking for the cheat. I checked whether the platform was harvesting credit through favorable attribution windows. I checked for selective reporting. I checked whether it was buying cheap inventory that happened to correlate with conversions it had nothing to do with causing.

Eventually I tested it against revenue lift—an outcome considerably harder to manufacture through convenient bookkeeping.

The performance held.

There was something real happening beneath the interface, and I wanted to know what it was.

Before the language existed

Artificial intelligence was a different category in 2011.

There was no common vocabulary separating artificial intelligence from machine learning, predictive modeling, statistical optimization, or automated decisioning. The terms collapsed into one another because the market lacked the fluency to keep them apart—which was, for certain vendors, the entire point.

The technology was also narrow. AI was not writing emails, generating images, or conversing with millions of people. Most consumers had no relationship with it whatsoever.

Programmatic advertising was one of the few commercial environments that could support machine intelligence at real scale. It produced enormous volumes of data. Millions of opportunities could be evaluated every second, each carrying signals about user, context, device, geography, time, historical behavior, creative, publisher, and economic value. The system could act, observe the result, and feed that result back into the next decision.

Queries per second created the conditions for machines to learn from staggering volumes of repeated bets.

A narrow manifestation of AI. But a real one.

Today most of the market has a chat window, and the complexity underneath it is invisible by design. I learned the old ways before the new interface arrived—the dark machinery beneath the magic: the data, thresholds, objectives, statistical assumptions, feedback loops, and human intervention required to make an intelligent system produce anything useful.

From evaluating the system to operating it

Rocket Fuel recruited me in 2012. I joined before the IPO as roughly its ninety-ninth employee and landed directly in the operational environment around the machine-learning platform.

I ran real-time bidding operations for a decision system evaluating millions of impressions per second—model testing, post-launch monitoring, campaign calibration, training-data integrity, and performance investigation across the bidding and optimization stack, with a small team and daily contact with product leaders and PhD-level engineers.

My job sat between the model and the market.

I needed to know how the quality of an input shaped an output. I needed to recognize when a model was learning from the wrong behavior, when a signal was being distorted, when a campaign simply lacked the data volume to conclude anything, and when the economic objective inside the system was quietly producing a result nobody wanted.

The most useful thing I learned in those years is that an optimization system will become spectacularly good at whatever you actually asked for, which is rarely what you meant.

A model can appear to be working while exploiting a flaw in the measurement framework around it. It is not lying. It is winning the game you built.

The work required practical fluency in logistic and linear regression, hierarchical modeling, Bayesian probability, feedback loops, thresholds, training-data volume, and the gap between statistical confidence and commercial action.

I did not design the mathematical architecture. I was the operator responsible for how the machine behaved once it met the world.

Two lessons survived everything since.

A model cannot rescue a poorly defined objective.

A model cannot overcome unreliable data.

The industry has since produced the evidence at scale. The RAND Corporation has reported that, by some estimates, more than 80 percent of AI projects fail. Gartner has pointed to poor data quality, inadequate risk controls, escalating costs, and unclear business value as recurring reasons generative-AI projects are abandoned after proof of concept.

The bottleneck was never the mathematics.

From model operations to market authority

Over the next five years my role expanded outward: client consulting, product strategy, sales strategy, product marketing, corporate strategy, analyst relations, executive communication.

I worked on folding the X+1 acquisition into a unified DSP and DMP platform—the unglamorous work of making two systems into one decisioning environment. Eventually I reported to Rocket Fuel's chief marketing officer as Director of Corporate Strategy and Innovation, responsible for helping the market understand what the technology was doing and why it mattered.

This was 2014 through 2017, when most executives had no working language for self-learning systems, predictive decisioning, or automated optimization.

The job was to translate an advanced technical system without reducing it to a slogan—to make it understandable while preserving the complexity that made it worth anything.

In 2014 I proposed the Marketing Activation Platform, built on three connected functions:

Integrate fragmented organizational data.

Activate it through machine intelligence.

Return what the system learned to the organization.

The argument was that marketing should operate as a continuously learning system. Data from CRM, media, customer interactions, identity, commerce, and business outcomes could be unified, acted on, and converted back into organizational knowledge. The system would run the campaign and, in the same motion, make the company smarter.

Much of that later arrived under other names: customer-data platforms, data clouds, journey orchestration, predictive audiences, AI-driven activation.

In 2017 I published a piece describing a future in which self-learning systems would process information, optimize workflows, and personalize experience at scale—separating signal from noise while combining machine intelligence with human judgment.

This period became an early form of AI commercialization: understanding a system well enough to explain it to customers, executives, analysts, and sales organizations without either mystifying or flattening it.

A technology does not become transformative because it works. It becomes transformative when people understand it, trust it, and reorganize themselves around what it makes possible.

Building the environment around the intelligence

After Rocket Fuel I consulted independently on machine-learning and data systems, often using TensorFlow to help organizations understand what their data could actually support.

Most of that work happened before any modeling.

Companies arrived with information scattered across flat files, disconnected databases, inconsistent identifiers, incompatible taxonomies, and systems built for reporting rather than learning. Before machine intelligence was worth discussing, the data had to become coherent: consistent identifiers, clear ownership, normalized fields, documented lineage, and an honest answer to whether the data being collected could support the decisions they wanted to automate.

The answer was frequently no. That answer saved several of them a great deal of money.

The visible model is a small part of an AI system. Collection, normalization, lineage, storage, permissions, processing, feature definition, measurement, deployment, monitoring, and feedback determine whether anything the model produces can be trusted.

A sophisticated model on top of disorganized infrastructure does not create intelligence. It creates sophisticated confusion.

Commercializing automated decision systems

At Adelphic and Viant I led product strategy, innovation, and customer success for the Adelphic demand-side platform.

The question was no longer whether automated decisioning worked. It was whether organizations could understand it, adopt it, operate it, and staff around it.

Customers needed to know which decisions the system was making, what it required, what the output meant, how to evaluate it, and where a human still belonged. Internally, people needed to sell the platform without exaggerating it, support it without undermining its automation, and build it without losing the customer's objective.

Part of my answer was Programmatic University—an operator certification curriculum and self-serve education layer built to teach the platform rather than staff around it. It cut support burden and accelerated enterprise adoption, and it hardened a conviction I still hold: an automated system scales at exactly the speed its users understand it.

I joined a distressed asset doing roughly $14 million in annual revenue. By the time I left, the platform was doing more than $300 million, with a 96 percent customer satisfaction rating sustained four years running. I did not do that alone—I ran the commercial motion around it, and served as a key lead on the IPO task force covering corporate narrative, investor relations, and roadshow positioning of the AI-platform story.

The technology mattered. The operating system around the technology made the growth possible.

AI at enterprise operating scale

At The Trade Desk I worked with machine intelligence at a different order of organizational scale.

Across several chapters I led a 130-person global enterprise services organization, built a net-new growth organization from zero, and carried growth accountability across a media portfolio exceeding $5 billion. I also served as a Product Senator—a formal role carrying roadmap authority and go-to-market sequencing for trading and optimization capabilities.

The Trade Desk was built around data-driven decisioning from the start. Machine learning informed media valuation, bidding, optimization, identity, measurement, and forecasting.

I did not build the core model. I connected the platform's intelligence to the organizations, operating practices, and commercial strategies required to get value out of it—including originating the automotive DSP initiative from concept validation into the roadmap, and helping carry a real-time sentiment tracker and an inventory quality index from client complaint to shipped capability.

I also worked on internal AI task-force efforts examining how agentic systems could improve platform navigation, analysis, optimization, and customer decision-making.

The central challenge never changed:

An AI feature is worth very little when it sits outside the way decisions actually get made.

Effective implementation requires knowing where the system acts, what it can access, when it should recommend, when it should execute, how it will be measured, and which decisions require a human being.

The model is one component. The system includes the workflow, the interface, the user, the incentives, the security boundary, the measurement framework, and the cost of being wrong.

From using AI systems to building with them

For the past two years I have been back in hands-on product development.

Agentic development environments have collapsed the distance between a business thesis and functioning software. I have used them to architect, design, build, test, secure, and ship applications, databases, websites, research environments, internal tools, design systems, workflow engines, and product proofs of concept.

This rests on almost two decades in graphic design, web development, advertising technology, data systems, product strategy, security, and commercial operations. I was building digital systems long before anything could generate code for me, and that foundation is doing more work than the model is.

Agentic development is most powerful when the person directing it understands what lies beneath the output: application architecture, data relationships, user behavior, security exposure, design hierarchy, commercial requirements, and the downstream cost of a decision made carelessly at two in the morning.

The work is sometimes called "vibe coding."

I understand the joke, and I would gently point out that the model has never once been the thing that got a product to market.

Producing code is easy now. Building a coherent, secure, useful, commercially viable system is exactly as hard as it was. The model accelerates implementation. It has no opinion about whether the thing should exist.

Building Patient Protect

The clearest expression of this is Patient Protect.

Working with a small founding team—my father Joseph, a security architect with more than thirty-six years in secure systems design and a background leading federal encrypted medical initiatives, and Angie, a registered dental hygienist and certified HIPAA consultant who has lived the compliance problem from inside an independent practice—I have built much of the company's product and commercial ecosystem through AI-assisted and agentic workflows.

That ecosystem includes a security-first HIPAA compliance platform, a public research institute, breach-intelligence systems, compliance assessments, risk models, workforce training, policy infrastructure, browser-security technology, a mobile application, and an AI compliance assistant.

These are production systems with real users, real permissions, real audit trails, real operating costs, and real reputational consequences.

The architectural claim I will defend is this: roughly twenty-five HIPAA requirements are satisfied by the platform's design rather than by a customer attesting that they have handled it.

That distinction is the whole thesis. Documentation-first compliance platforms are built to survive an audit. This one was built to prevent the breach that triggers the audit. Everything else—endpoint counts, component counts, environment counts—is scope, not quality. A bad application can have two hundred endpoints. So can a good one.

Patient Protect Signal is a shipped iOS application bringing breach intelligence, compliance assessments, risk tools, and research into a mobile environment. Free, no subscription, no protected health information collected.

HIPAA Shield is a client-side browser extension that detects protected health information before it is typed or pasted into a consumer AI system—Social Security numbers, dates of birth, medical record numbers, validated card numbers, diagnosis codes, clinical terminology. Detection runs locally. The extension makes zero network requests and transmits no telemetry. We released the full source under the MIT license so any practice can install it, audit it, or fork it, customer or not.

We built it because the exposure is already ordinary. Netskope's 2025 healthcare report found that 88 percent of healthcare organizations were using cloud-based generative AI apps directly, while regulated healthcare data was the dominant data-policy violation category across personal apps, generative AI apps, and other unapproved destinations. Wolters Kluwer Health survey research separately found that 57 percent of healthcare professionals had encountered or used unauthorized AI tools in their organizations, with 17 percent saying they had used one themselves.

Consumer AI tiers generally do not create HIPAA coverage by default. OpenAI's current HIPAA-eligible product list says coverage requires a Business Associate Agreement and is limited to products such as ChatGPT for Healthcare, ChatGPT for Enterprise with Regulated Workspace, ChatGPT FedRAMP, ChatGPT for Clinicians, API with Modified Retention, and API FedRAMP with Modified Retention. When protected health information goes into an uncovered tool, the covered entity may have made an impermissible disclosure—which 45 CFR §164.402 presumes to be a breach unless the entity can demonstrate low probability of compromise through a documented risk assessment, or unless a specific exception applies. 45 CFR §164.404 then governs individual notification after discovery of a breach.

That is a meaningfully worse position than most practices believe they are in, and it is not the same thing as a guaranteed notification event. The distinction matters, and getting it wrong in either direction is how compliance vendors lose credibility.

What it looks like from the inside is simpler: the technology arrived years ahead of the governance, and the people absorbing the risk are the least equipped to see it coming.

PIPAA and private intelligence

PIPAA is the most direct expression of my current AI work: a domain-specific HIPAA compliance agent built around regulatory knowledge, platform awareness, user permissions, workflow execution, and safeguards for protected health information.

It understands a practice's compliance state, cites federal regulation by provision, identifies open risk, creates tasks in the compliance queue, and supports documentation. Ask whether adding a vendor triggers a risk-assessment update and it returns the governing citation, then makes the task. A public version is available at Ask PIPAA.

The architecture reflects where it operates. The production design does not send prompts or protected health information to an uncontrolled third-party model. Development runs on dedicated local hardware with a Llama-based inference environment; an on-site deployment serves organizations requiring zero network exit. A redaction layer strips protected health information from prompts before processing—the architecture was built so this should never fire, and we built it anyway.

The feature I am proudest of is that it refuses.

When PIPAA lacks sufficient context, it says so instead of guessing. In a regulated domain, a confident wrong citation is worse than no answer at all, and most of the engineering effort in that system went into teaching it where its authority ends. That is a strange thing to optimize for. It is also the only version of the product that deserves to exist.

The model is one part of the architecture. The rest is regulatory grounding, platform context, permissions, data access, redaction, local inference, task creation, auditability, workflow integration, security boundaries, and human accountability.

A useful agent needs more than language. It needs to know where it is, what it may touch, what it may change, what evidence supports its answer, and when to get a person.

The lesson, relearned

In 2026 I nearly published a number that was wrong by orders of magnitude.

Through the Secure Care Research Institute we produce quarterly breach research from federal disclosure data, state attorney general filings, and enforcement records. Preparing a report, we were building on an aggregated public healthcare breach dataset carrying a headline figure in the billions of affected individuals.

Nothing about the file looked wrong. It was well-formatted, openly licensed, widely used, and confidently labeled.

Two things were happening inside it.

The same incidents had been ingested through multiple source paths and never deduplicated, so major breaches were being counted repeatedly. And the file contained AI-modeled rows—synthetic estimates generated to fill reporting gaps—sitting in the same columns as verified disclosures, with nothing in the schema to distinguish an observation from an inference.

We rebuilt the methodology. The full quarterly was replaced with a verified brief covering a smaller number of individually confirmed disclosures. Modeled rows were stripped from the public download. A claim we had planned to make about attack-vector classification was cut rather than supported by data I could not stand behind.

The corrected work is smaller, slower, and considerably less impressive on a chart.

It is also true.

This is the 2012 lesson wearing a new coat. Then, a model could look like it was working while quietly exploiting a flaw in the measurement around it. Now a dataset can look authoritative while mixing synthetic output into an evidentiary record—and the next system trained on that record inherits the error as fact, laundered through one more layer of citation.

Generative tools have made this failure faster, cheaper, and much harder to see. The synthetic and the observed no longer look different.

If I have one habit worth transferring from fifteen years of this, it is asking where a number came from before deciding what it means.

The framework beneath the work

In 2025 I published a framework on SSRN called The Infrastructure of Trust.

Its argument is that as intelligence commoditizes, the binding constraint on AI systems will not be computational capacity but trust capacity—whether the inputs, incentives, and permissions surrounding a model can be verified at all.

None of the underlying components are new. Data provenance, incentive alignment, consent, security, transparency, and data sovereignty are all mature fields with their own literatures. The paper's contribution is synthesis: pulling fragmented streams from cybersecurity, data governance, behavioral economics, and responsible AI into a single dependency structure, and arguing that the layers constrain one another in a specific order.

It also explains why my work looks scattered and is not.

Patient Protect is the integrity layer, applied to healthcare data. Ad Certainty—a publisher platform built around verified consent, auditable trust signals, and publisher economics—is the consent layer, applied to media. RADIX, my philosophical work, is the coherence question, applied to people.

Enormous capital is moving into artificial intelligence. Very little of it is moving into the infrastructure required to verify what those systems are being fed.

That asymmetry is the whole of my professional interest.

The human work ahead

My work has always lived around the intelligence rather than inside it: between the model and the market, the code and the customer, technical possibility and human consequence.

That vantage produced real respect for what these systems can do, and removed any temptation to mythologize them.

AI fails. It breaks. It hallucinates. It misreads intent, amplifies weak assumptions, and executes flawed instructions with total confidence and excellent grammar.

It carries the fingerprints of the people who designed it, the data used to train it, the objectives it was given, and the incentives of whoever deployed it. It is an extension of human design, and human design has always been imperfect.

Garbage in still produces garbage out. AI simply lets the garbage travel faster, farther, and with a much better interface.

This is why integrity matters more as intelligence scales, not less. A system acting across millions of decisions needs clear rules, real boundaries, reliable evidence, security controls, and people with the instinct to recognize when an answer is internally coherent and completely wrong.

Scale does not remove the need for judgment. It magnifies the cost of its absence.

The highest use of these tools is the expansion of human capacity. They can compress research, absorb repetitive work, surface patterns, accelerate experimentation, and shorten the distance between an idea and its first working version. They can give one person the reach that used to require a department.

That leverage should buy back room for imagination, craft, care, and strategy.

I have limited interest in a future where AI is celebrated mainly as a way to reduce operating expense. A technology this consequential deserves a larger ambition than headcount reduction. Its real promise is helping people see farther, build faster, and spend more of their working lives on things that require judgment, taste, courage, and invention.

It should augment the artist, the engineer, the doctor, the operator, the researcher, the founder—removing friction around a person's gift without erasing the person responsible for it.

The cost of effortless creation

Creation is losing its friction fast.

Software in minutes. Images in seconds. Strategies, articles, decks, campaigns, entire brand identities, arriving fully formed and slightly damp.

This democratizes making, which is good, and introduces a cultural risk that is not.

When everyone draws from the same models, writes the same prompts, adopts the same optimization patterns, and accepts the first polished answer, everything starts to sound the same. The language becomes fluent and bloodless. The design becomes beautiful and familiar. The idea arrives complete and belongs to no one.

The next real advantage will be discernment.

Taste will matter. Point of view will matter. Lived experience will matter. The strange instinct that rejects the statistically obvious answer will matter enormously, along with the patience to push past the model's first approximation of what "good" is supposed to look like.

Human work has always carried irregularity—obsession, contradiction, memory, humor, private symbolism, and choices no performance data would ever justify. Those are the parts worth defending.

AI can assist the making. Meaning still has to come from the maker.

Foundations before fashion

Mastering the new tools requires understanding what sits underneath them.

An application without data integrity is a failure that simply has not happened yet. Data without security is exposure. Secure infrastructure without a useful product is an engineering exercise. A useful product without distribution, positioning, and adoption is an artifact nobody encounters.

Every layer matters. The interface gets the attention; the full system determines whether anything lasts.

That is how I build. The data, the permissions, the security boundary, the model behavior, the user experience, the economics, the narrative, the distribution, and the human outcome are parts of one organism, and I have never been able to think about them separately.

The new tools make construction faster. They do not make the foundations optional.

What I want to build is straightforward enough to say plainly: systems that expand human agency, protect sensitive information, preserve authorship, and hand people back the parts of their work that make them worth hiring in the first place. I want intelligence to produce more possibility rather than smaller organizations—more artists, more builders, more independent thinkers, more people capable of bringing something into the world that would not have existed otherwise.

AI will scale whatever we hand it: our judgment, our incentives, our imagination, and our flaws.

That is the argument for handing it something worth scaling.