The other day I noticed a taxonomy used on one of the NoSQL Database blogs that went like this:
Types of NoSQL systems
- Core NoSQL Systems
- Wide column stores
- Document stores
- Key-value / tuple stores
- Eventually consistent key-value stores
- Graph databases
- Soft NoSQL Systems (not the original intention …)
- Object databases
- Grid database solutions
- XML databases
- Other NoSQL-related databases
I, perhaps obviously, take some umbrage at having MarkLogic (acceptably classified as an XML database) being declared “soft NoSQL.” In this post I’ll explain why.
Who decided that being open source was a requirement to be real NoSQL system? More importantly, who gets to decide? NoSQL – like the Tea Party – is a grass-roots, effectively leaderless movement towards relational database alternatives. Anyone arguing original intent of the founders is misguided because there is no small group of clearly identified founders to ask. In reality, all you can correctly argue is what you think was the intent of the initial NoSQL developers and early adopters, or — perhaps more customarily — why you were drawn to them yourself, disguised or confused as original founder intent.
As mentioned here, movements often appear homogeneous when they are indeed heterogeneous. What looks like a long line of demonstrators protesting a single cause is in fact a rugby scrum of different groups pushing in only generally aligned directions. For example, for each of the following potential motivations, I am certain that I can find some set of NoSQL advocates that are motivated by it:
- Anger at Oracle’s heavy-handed licensing policies
- The need to store unstructured or semi-structured data that doesn’t fit well into relations
- The impedance mismatch with relational databases
- A need and/or desire to use open source
- An attempt to reduce total cost
- A desire to land at a different point in the Brewer CAP Theorem triangle of consistency, availability, and partition tolerance
- Coolness / wannabe-ism, as in, I want to be like Google or Facebook
(Since this was a source of confusion in prior posts, note that this is not to claim the inverse: that all NoSQL advocates are motivated by all of the possible motivations.)
I’d like to advocate a simple idea: that NoSQL means NoSQL. That a NoSQL system is defined as:
A structured storage system that is not based on relational database technology and does not use SQL as its primary query language
In short, my proposed definition means that NoSQL (broadly) = NoSQL (literally) + NoRelational. In short: relational database alternatives. It does not mean:
- NoDBMS. We should not take NoSQL to exclude systems we would traditionally define as DBMSs. For example, supporting ACID transactions or supporting a non-SQL query language (e.g., XQuery) should not be exclusion criteria for NoSQL.
- NoCommercialSoftware. While many of the flagship NoSQL projects (e.g., Hadoop, CouchDB) are open source projects, that should be not a defining criterion. NoSQL should be a technological, not a delivery- or business-model, classification. Technology and delivery model are orthogonal dimensions. We should be able to speak of traditionally licensed, open source licensed, and cloud-hosted NoSQL systems if for no other reason than understanding the nuances of the various business/delivery models is a major task unto itself. Do you mean open source or open core? Is it open source or faux-pen source? Under which open source license? How should I think of a hosted subscription service that is a based on or a derivative of an open source project?
Recently, I’ve heard a piece of backpeddling that I’ve found rather irritating: that NoSQL was never intended to mean “no SQL,” it was actually intended to mean “not only SQL.” Frankly, this strikes me as hogwash: uh oh, I’m afraid that people are seeing us as disruptors and it’s probably easier to penetrate the enterprise as complementary, not competitive, so let’s turn what was a direct assault into a flanking attack.
To me, it’s simple: NoSQL means NoSQL. No SQL query language and no relational database management system. Yes, it’s disruptive and — by some measures — “crazy talk” but no, we shouldn’t hide because there are lots of perfectly valid (and now socially acceptable) reasons to want to differ from the relational status quo.
In effect, my definition of NoSQL is relational database alternative. Such options include both alternative databases (e.g., MarkLogic) and database alternatives (e.g., key/value stores). This, of course, then cuts at your definition of database management system where I (for now at least) still require the support of a query language and the option to have ACID transactions.
By the way, I understand the desire to exclude various bandwagon-jumpers from the NoSQL cause. Like most, I have no interest in including thrice-reborn object databases in the discussion, but if the cost of excluding them is excluding systems like MarkLogic then I think that cost is too high. Many people contemplating the top-of-mind NoSQL systems (e.g., Hadoop) could be better served using MarkLogic which addresses many typical NoSQL concerns, including:
- Vast scale
- High performance
- Highly parallel shared-nothing clusters
- Support for unstructured and semi-structured data
All with all the pros (and cons) of being a commercial software package and without requiring reduced consistency: losing a few Tweets won’t kill Twitter, but losing a few articles, records, or individuals might well kill a patient, bank, or counter-terrorism agency. BASE is fine for some; many others still need ACID. Michael Stonebraker has some further points on this idea in this CACM post.
I’d like to suggest that we should combine the ideas in this post with the ideas in my prior one, Classifying Database Management Systems. That post says the correct way to classify DBMSs is by their native modeling element (e.g., table, class, hypercube). This post says that NoSQL is semi-orthogonal – i.e., I can imagine a table-oriented database that doesn’t use SQL as its query language, but I doubt that any exist. Applying my various rules, the combined posts say that:
- Aster is a SQL database optimized for analytics on big data
- MarkLogic is an XML [document] database optimized for large quantities of semi-structured information and a NoSQL system
- CouchDB is a document database and a NoSQL system
- Reddis is a key/value store and a NoSQL system
- VoltDB is a SQL database optimized to solve one of the two core problems that NoSQL systems are built for (i.e., high-volume simple processing)
Finally, I’d conclude that even with these rules I have trouble classifying MarkLogic because of multiple inheritance: MarkLogic is both a document database and an XML database, it is difficult to pick one over the other, and I there certainly are non-document-oriented XML database systems. Similar issues exist with classifying the various hybrids of document databases and key/value stores. So while I may have more work to do on building an overall taxonomy, I am absolutely sure about one thing: MarkLogic is a NoSQL system.
—
* The “Yes, Virginia” phrase comes from a 1897 story in the New York Sun. For more, see here.