Different Types of Replication
To get right to the point, there are three fundamental forms of data replication. For example, disk blocks can be replicated (block-based replication, such as in Redundant Arrays of Inexpensive Disks— RAID), records can be replicated (in databases), or entire files can be replicated (such as AWS S3 buckets, Google Cloud Storage, or Microsoft Azure Blob storage). All have their strengths and weaknesses and, therefore, specific areas of application. In “Cloud Speak,” we refer to these as“Patterns and Anti-Patterns.” Here, the “Patterns” are the correctly applied replication types, and the “Anti-Patterns” are incorrectly applied replication types in a specific scenario.
Block-based storage
Block-based storage is commonly found in the world of Storage Area Networks (SANs). This ensures that data is distributed across multiple disks, so that the failure of a single disk does not result in data loss. It is a highly efficient form of data replication because only changed blocks are replicated, as opposed to replicating the entire file. This approach is not limited to SANs; there is also open-source software available for block-based replication (including over a network), such as DRBD.
Block-based replication is typically used to achieve redundancy (of disks, for example). When replication is not between local disks but to a different disk at another location, the goal is not only disk redundancy but also data availability. The (physical) location of the data is then eliminated as a Single Point of Failure (SPoF). However, even such a solution has its limitations. If data is deleted on one side, it is also deleted on the other side. Therefore, it does not provide protection against that scenario. It is not disaster recovery (DR). An off-site backup, for example, is.
Record-based replication
Record-based replication is used with databases. Databases handle read and write operations on records, for example. In replication, only write operations are replicated, since read operations do not change anything. The write operations are performed again on the other side so that both databases are once again in sync. Often, the replicated database is a “read-only” database (also called a “slave”). The replicated (slave) database can, for example, be used effectively for backups. This has no performance impact on the “primary” or “master” database.
Replication is therefore also used here to increase availability (and does not involve disaster recovery). The backup created from the “Slave” database is, however, a DR measure. The replicated write operations can be stored in a redo log file (Oracle technology). These redo log files are then often included in the backup as well. By combining the database backup with the redo log files, a“point-in-time restore”can be performed.
Databases offer countless options for replication. For example, MySQL and MariaDB provide extensive capabilities for this purpose. In the example above, the database replica was not used for production. It is used for backup purposes and has no end users. However, both databases (“Master” and “Slave”) can be used for read operations, as long as changes are written only to the “Master.” In general, a database has 10 times as many read operations as write operations. A “Master” with 9 “Slaves” can therefore be a way to increase the overall performance of the entire system. But even though performance is higher, the “Master” is still a Single Point of Failure (SPoF). One of the “Slave” databases can be promoted to “Master,” but that takes time and therefore results in downtime for write operations.
“Master”—“Master” setup
ACID
A term used in databases is ACID (Atomicity, Consistency, Isolation, and Durability). These are characteristics of a database that ensure the accuracy and completeness of data. The “Atomicity” property states that database operations are indivisible and are therefore either executed and saved correctly or not at all. It prevents operations from being partially executed and ensures that data is neither missing nor corrupted.
Another option is to set up a “Master-Master” configuration. In this setup, both “Master” databases can receive write operations and replicate them to the other “Master.” It is essential to prevent records with the same record number from being written. Therefore, one “Master” is designated to handle “even” inserts, and the other “Master” handles “odd” inserts. Both databases can be used with local performance—even when they are located far apart. This system can be expanded as needed, for example, with a 3- “Master” setup (see the figure). This increases both data availability and performance. However, a “Drop Table” command is replicated to all “Masters,” so this is not a disaster recovery (DR) solution.
One final note about the “Master” multi-database setup. When the databases are located far apart, replication latency increases. Both databases can be used locally with excellent performance. This is a good setup for active-active (or “Multi-Region”) designs. In this scenario, replication may “lag behind” (consider the RPO; see the section on RPO and RTO).
File-based replication
In cloud environments, the (simple) replication of files is a basic feature. This existed even before the cloud (Ceph, for example ), but the cloud has certainly made it popular. These days, everyone knows what an AWS S3 bucket is. With file-based replication, files are stored on multiple storage systems (and often in multiple locations). As soon as a file changes, a completely new version of the file (or object) is created—a process also known as“Copy-on-Write.”As a result, when a file (or object) changes, all copies must be updated immediately. This takes time, which is why it’s often referred to as“Eventually Consistent”storage. The file (or object) can be read from any location (once it has become consistent). This results in high performance. Version control is also (relatively) simple; after a “Copy-on-Write,” the old version is not deleted.
A drawback of the “Copy-on-Write” principle is that if even a single bit of a file is changed, the entire file must be rewritten. Storing a database file on an object store (such as S3) is therefore not efficient. Every record change results in a new copy of the entire database file. This leads to poor performance; it is what is known as an“anti-pattern.”It can be argued that object storage (file-based replication) thus improves both performance (in the right situations, the “patterns”) and availability. It can also provide disaster recovery (DR), provided that version control is enabled. A final advantage of this type of storage technology is that it offers virtually unlimited scalability in terms of storage capacity. It is exceptionally powerful in its simplicity.
But can an object storage file system be mounted?
Block-based file systems can be mounted “on” or “to” a computer. They are thus treated as local storage, even though they are not. This is not common with object storage. It has an Application Programming Interface (API) that can be used. Read, write, create, and delete operations can be performed via this API. However, if you really want to, it is possible to mount an object store. The performance of such a solution is low, mainly due to the “Copy-on-Write” principle. The object store behaves fundamentally differently from local storage, and as a result, it has few use cases (patterns).
RTO and RPO
When designing an environment, the RTO and RPO that must be achieved are determined. The Recovery Point Objective (RPO) is the maximum amount of data loss that an environment is allowed to experience. The Recovery Time Objective (RTO) is essentially the maximum downtime, but it can also be viewed as the environment’s availability. These values are often determined during a Business Impact Analysis (BIA).
The RTO is achieved through all the measures taken to guarantee or restore availability. The RPO is primarily determined by the latency of data replication and backup/restore (DR). But if a network outage occurs before data has been fully replicated, how much data has been lost? How much data loss is actually acceptable? RPO is therefore a complex topic, because how do you know which data has been replicated during a failure and which has not? How do you determine which of the two sides contains the correct data if they are no longer in sync? This is also referred to as the data “GAP.” This brings us to the topic of synchronous and asynchronous data replication. More on this in the “Timing” section.
Which data is now correct after we experienced a replication failure?
With a database, the “timestamp” can be used. The last record written correctly is the last valid data; everything after that is incorrect. A prerequisite for this is that both sides have the correct time. With block-based replication, a technique called “fencing” is used. “Fencing” is used to determine which (storage) side contains the truth.
What about backups?
You might—rightly—wonder whether backup/restore is a data replication technique. That’s certainly worth discussing. In most cases, however, backup/restore is not considered a replication technique. Backup/restore is more about retention and recovery. It creates a “point-in-time” copy of data (including errors). Data replication, on the other hand, is a technique used to ensure“business continuity.”What adds to the confusion is that object storage can also be used to store backups. It offers high availability and supports version control. This debate is difficult to settle. It is therefore crucial to answer the question, “What needs to be achieved?”
What about off-site backups?
Object Store offers good options for storing backups. Not only does it support version control, but it also allows you to use other regions. This brings us close to the concept of off-site backups.
Granularity
The three replication technologies all have their advantages and disadvantages, their “Patterns & Anti-Patterns.” One major difference between these techniques, however, is their granularity. The smallest replication unit is block-based replication, which involves, for example, 4-kilobyte blocks. Records are usually much larger, and files are often many times larger still.
Topology
The article previously discussed “one-way” replication and “two-way” replication. This is also referred to as the replication topology.
Deduplication
When file storage is analyzed, it becomes apparent that the same files are often stored multiple times. That’s a shame, because replication already keeps multiple copies. Copies of copies are unnecessary and costly. Some storage systems remove these duplicates without the end user even noticing. This is called deduplication. Deduplication saves storage capacity and, therefore, costs.
Compression
In addition to deduplication, data and files can also be compressed. This involves reducing the size of files through compression, thereby saving storage capacity. The greatest benefit is achieved when both deduplication and compression are combined. However, both processes require computing power (overhead) and therefore introduce latency (albeit theoretically).
Encryption
One topic that comes up regularly is encryption. Most cloud providers and hyperscalers offer this“out of the box.”It can be applied to local (block) storage, databases, and object storage. It’s a good idea to use encryption—it’s easy and costs virtually nothing. However, it doesn’t protect you from the cloud provider or hyperscaler itself. They still have access to the encryption keys. If this needs to be prevented, a company or organization will have to manage its own encryption keys. This is feasible, but it doesn’t integrate very well into the cloud environment—for example, using a Hardware Security Module (HSM). This does, however, provide protection against the cloud provider or hyperscaler, as well as against information requests from governments and intelligence agencies. This is therefore an important aspect of the discussion surrounding data sovereignty. This encryption of stored data is also referred to as“Data at Rest”encryption. In addition to “Data at Rest” encryption, there is also“Data in Transit”encryption, which applies to network traffic. Both are forms of encryption, and both protect against different risks.
Caching
One topic mentioned here is caching. Caching—in general—is the approach of storing frequently requested data on fast disk (SSD), in memory, or in a geographically close location. This allows the data to be delivered very quickly. Operating systems such as Linux, OS X, and Windows also use this concept. Databases also utilize caching. This enables frequently requested records to be retrieved quickly. For example, this is useful for session information that is frequently requested. Caching is a topic related to data replication because, on the one hand, it can speed up replication, and on the other hand, it can disrupt replication. We won’t go into further detail on this topic here, but you’ll explore concepts such as “cache hits,” “cache misses,” and “flushing the cache.”
Timing
Another interesting aspect of data replication needs to be discussed in more detail:latency. This is important in replication. It was mentioned earlier in the context of RTO and RPO. The physical distance between the storage systems we are replicating determines how long it takes to replicate that data. The data is not secure until it has been fully written to both sides. When the distance is great, this begins to play a role. With synchronous replication, confirmation is only given once both sides have written the data. Due to the distance, an end user must then wait longer. At very long distances, this can become a problem. In such cases, it’s better to opt for asynchronous replication. With this method, the end user can continue working once the data has been written locally. The data is then replicated later. This is good news for performance, but less so for data security. After all, it’s possible that the data hasn’t been fully replicated yet when a disruption occurs. The storage is then no longer the same. Which side is correct if bidirectional replication was used? There are, of course, solutions for this, but they fall outside the scope of this article.
Data Sovereignty
And why was all of this so important for data sovereignty? In light of data sovereignty, companies and organizations might consider storing data in multiple locations—for example, to ensure the (geo)location of the data (data residency). Various techniques also offer the possibility of data recovery (DR). This is something that certainly impacts data sovereignty. The techniques described are therefore important for achieving such objectives. And because no single technique solves all problems and situations differ from one another, it is useful to be familiar with all techniques and to “master” them. Perhaps a combination of techniques can—or even must—be used.
Although it is not a replication technology, open-source software (OSS) deserves a mention here. Because data sovereignty is about gaining “control” over data, the software used has a major impact on that. The best way to maintain (or regain) control is to retain control over the software being used. OSS is therefore the best choice one can make. An interesting detail, by the way, is that all cloud providers and hyperscalers use OSS. Sometimes they go so far as to rename OSS and market it as their own product. This is, incidentally, legal. Even if the renamed version is used, one can always revert to the OSS version, ensuring that data remains accessible.
Another advantage of using OSS is that it is the best way to ensure the use of open standards. OSS has no choice but to use open standards, and that guarantees access to the data.
In summary, it can be said that understanding “Patterns and Anti-Patterns” and mastering the technology of data replication are most likely essential to achieving data sovereignty. Data sovereignty involves mitigating geopolitical risks associated with data storage. These techniques will help you achieve that.