The argument Codd made

In June 1970, Edgar F. Codd published "A Relational Model of Data for Large Shared Data Banks" in the Communications of the ACM. He was a mathematician working at IBM's research division in San Jose, and the paper was a direct attack on the dominant database architectures of the day: hierarchical systems like IBM's own IMS, and network databases organized according to the CODASYL specification. Both required the application to navigate physical storage structures explicitly — to know not just what data it wanted, but where that data lived and how to traverse the pointers to reach it. Codd argued that this was the wrong abstraction entirely.

His alternative rested on a branch of mathematics called predicate logic and set theory. Data should be organized into relations — what we now call tables — where each row is a tuple and each column is an attribute with a defined domain. The structure should be entirely logical, not physical. A user or a program should be able to ask for data by describing what they want, leaving the system to decide how to retrieve it. This separation of logical structure from physical storage, which Codd called data independence, was the intellectual core of the paper. The applications that ran on top would not break just because the underlying storage changed.

A bound ledger open under one lamp, black surround

Before the software, the deadline. The close is older than every system that serves it.

Thraex picture desk

That sounds obvious now. In 1970 it was a minority position inside a company whose flagship database product would have been disrupted by it. IBM filed the paper and moved slowly, continuing to invest in IMS. The relational model would find its first serious commercial exploiter elsewhere.

From theory to product, and the vendor who moved fastest

Larry Ellison read about Codd's work — by some accounts the specific impetus was a 1976 paper by IBM researchers describing a prototype called System R, which was IBM's own attempt to build a relational database engine. Ellison, working with Bob Miner and Ed Oates, built a commercial relational database and deliberately named it Oracle, the same codename as a CIA project Ellison, Miner and Oates had worked on at Ampex. Oracle version 2 shipped in 1979, before IBM had released a commercial relational product of its own. That head start hardened into a durable market position that still shapes enterprise data infrastructure today.

Chronology

  1. 1970Codd publishes the relational model paper in Communications of the ACM
  2. 1976IBM publishes the System R paper; Ellison cites it as a prompt
  3. 1979Oracle version 2 ships, the first commercial relational database
  4. 1985Codd publishes his Twelve Rules in Computerworld
  5. 1986–1987SQL standardized by ANSI and ISO
  6. 1992SAP R/3 ships, organizing its data model on relational tables
  7. Early 2000sNoSQL movement emerges for non-structured workloads

The query language that System R had developed — Structured Query Language, or SQL — became the interface through which virtually every relational database would be addressed. SQL was standardized by ANSI in 1986 and by ISO the following year, which meant that the syntax a developer learned on one system was portable, in principle, to another. In practice the vendors all extended SQL with proprietary features, stored-procedure dialects, and optimizer hints that created friction between systems. The SQL standard itself has been revised repeatedly — SQL-92, SQL:1999, SQL:2003, and beyond — each revision adding features that the vendors had already implemented in incompatible ways.

Codd watched what was done with his model and was not entirely satisfied. In 1985 he published twelve rules — quickly called Codd's Twelve Rules — that he said a database must satisfy to be described as truly relational. The exercise was partly polemical: several commercial databases were marketing themselves as relational while violating properties the model required. Few systems have ever satisfied all twelve rules in full. The rules survive as a reference point for what the model actually demands, not as a certification anyone receives.

A stack of unlabelled floppy disks, macro

Distribution in the 1980s: one program, one machine, one box.

Thraex picture desk

Why it won, and what winning cost

The relational model succeeded for reasons that were partly technical and partly structural. Mathematically, it gave database design a discipline: normalization, the process of organizing a schema to eliminate redundant data and ensure logical dependencies are properly structured, could be applied rigorously. First normal form, second, third, Boyce-Codd — these were provable properties, not conventions. A schema that satisfied them had documented guarantees about how updates would propagate without corruption.

But the relational model also succeeded because it arrived at the right commercial moment. The 1970s and 1980s saw enterprises accumulating structured data — orders, inventory, accounts, payroll — at a rate that the older hierarchical systems struggled to serve flexibly. SQL gave business analysts a language that was at least legible without a programmer intermediary. The query model fit the kinds of questions finance and operations departments actually needed to ask.

Key concepts

Lifted from the piece
Relational model
Codd's mathematical framework: data as tables of tuples, queried by logical description rather than physical navigation
Data independence
the separation of logical schema from physical storage; applications survive storage changes
Normalization
the process of structuring a schema to provable rules (normal forms) that eliminate redundancy and prevent update anomalies
SQL
the query language developed in IBM's System R project; ANSI/ISO standard from 1986
Codd's Twelve Rules
1985 criteria for what truly constitutes a relational system; more polemical checklist than certification

What winning meant, in practice, was that the table-row-column structure became the default assumption baked into every layer of enterprise software above the database. SAP R/3, which shipped in 1992, organized its entire data model around relational tables — tens of thousands of them. Oracle's own application suite extended from the database engine upward into ERP, CRM, and HR. When Workday built its cloud-native HR and finance system in the mid-2000s it still stored its data relationally, even as it abstracted the schema from customers. Salesforce, whose multi-tenant architecture separates customer data within shared infrastructure, runs on Oracle databases at its foundation.

The consequence is a form of structural lock-in that operates at a level beneath any individual vendor. A customer who migrates from one ERP to another is not just moving records between systems; they are translating between two relational schemas, each of which has encoded decades of decisions about how the business's data is shaped. The data migration project that follows is often the most expensive and risky part of any such move, and its difficulty is precisely because both source and destination speak SQL but organize what they say in entirely different ways.

What the model did not anticipate

Codd's paper assumed structured data: defined attributes, defined domains, relations whose shape was known in advance. The document stores, graph databases, and wide-column stores that emerged from the early 2000s onward were responses to data that did not fit this assumption — web-scale event logs, social graphs, content with no fixed schema. These systems were grouped under the label NoSQL, which was more a gesture of departure than a coherent alternative model.

The departure has been partial. Most organizations running NoSQL systems for specific workloads continue to run relational databases for their financial records, their HR data, and anything subject to audit requirements. The properties that make relational systems difficult — the rigid schema, the transactional guarantees, the requirement to model data before you store it — are the same properties that make them trustworthy enough to close the books on. Fifty years after the paper, the schema Codd described is still the one an accountant's records live in.