Safely replacing CSV files in Python: Preserving original data even if write fails mid-process

プログラミング・Web開発カテゴリを表すパンダのイラスト Programming / Web Development

About this article
This article was generated using an AI-assisted automated workflow. Referring to the official Python documentation for csv, tempfile, and os.replace, we created code that preserves the original CSV in the event of an intermediate error. Sample execution, syntax checks, and three unit tests were performed on Python 3.13.5.

Verification Status: 🧪 Executed locally on Python 3.13.5 with 3 successful unit tests.
Integration behavior with WordPress or enterprise systems, Windows environments, and sudden power outages have not been verified.

When updating a CSV file in Python, writing directly to the existing file may leave a corrupted or partially written file if an exception occurs during the process.

Bywriting new content completely to a temporary file before replacing it,you can easily prevent situations where the original file is partially overwritten due to standard write exceptions.

Here, we test a program that uses only standard libraries to replace the file with the new CSV upon successful completion, while preserving the original CSV if validation errors occur. Since the destination is a temporary directory created for learning purposes each time, enterprise data is not affected.

How does it differ from writing directly?

If you open an existing CSV in the following manner, the file contents are truncated when writing begins.

# 注意: 既存のファイルを直接書き換える例

with open("batch.csv", "w", encoding="utf-8", newline="") as fp:
    fp.write("id,status\n")

    # ここで処理が停止すると、元データは戻りません。

If you want to "switch to new content only when writing succeeds," the approach of creating a separate file first is effective.

ProcessDirect overwriteUsing a temporary file
Write startWrite to original fileWrite to temporary file
Validation error during the processThe original file may be modified midwayThe original file can be preserved prior to replacement
Upon completionExit normallyReplace using os.replace
Cautionary notesPartial writesAdditional measures for concurrent updates or power outages are required separately

The operation of replacing the filename upon success of os.replace is defined as an atomic operation in POSIX environments. However,the fact that the filename replacement is atomic is a separate matter from ensuring that the new data always persists after a sudden power outage. The official documentation isPython os.replace.

Verified sample code

This code targets Python 3.10 and later. It uses the standard libraries csv, os, tempfile, and pathlib.

from __future__ import annotations

import csv
import os
import tempfile
from pathlib import Path
from collections.abc import Iterable, Mapping

COLUMNS = ("id", "status")


def replace_csv(target: Path, rows: Iterable[Mapping[str, str]]) -> None:
    """CSVを書き終えるまで元ファイルを変更しない。"""
    target = Path(target)
    temporary: Path | None = None
    try:

        # 置換先と同じフォルダーに一時ファイルを作成する

        with tempfile.NamedTemporaryFile(
            mode="w", encoding="utf-8", newline="", delete=False,
            dir=target.parent, prefix=f".{target.name}.", suffix=".tmp"
        ) as fp:
            temporary = Path(fp.name)
            writer = csv.DictWriter(fp, fieldnames=COLUMNS, extrasaction="raise")
            writer.writeheader()

            for row in rows:

                # 入力列の過不足も検出し、失敗したら置換しない

                if set(row) != set(COLUMNS):
                    raise ValueError("CSV row must have exactly id and status")
                writer.writerow(row)

            # Pythonのバッファを出し、OSにファイル内容の同期を依頼

            fp.flush()
            os.fsync(fp.fileno())

        # 全件を書き終えてから、元CSVの名前へ置き換える

        os.replace(temporary, target)
        temporary = None
    finally:

        # 置換前に失敗したら、作った一時ファイルだけを削除

        if temporary is not None:
            temporary.unlink(missing_ok=True)


if __name__ == "__main__":

    # 毎回新しいディレクトリにサンプルCSVを作る

    with tempfile.TemporaryDirectory(prefix="csv-atomic-demo-") as folder:
        csv_path = Path(folder) / "batch.csv"
        csv_path.write_text("id,status\nold,OLD\n", encoding="utf-8")

        replace_csv(csv_path, (
            {"id": "A-001", "status": "DONE"},
            {"id": "A-002", "status": "PENDING"},
        ))

        with csv_path.open(newline="", encoding="utf-8") as fp:
            entries = list(csv.DictReader(fp))
        print(f"[RESULT] rows={len(entries)} ids={[x['id'] for x in entries]}")

The executable file hosted on GitHub contains the same content.

You can run the following command from the terminal in the sample folder.

python3 atomic_csv.py
python3 -m unittest -v test_atomic_csv.py

Actual verified output

The test environment for Python 3.13.5 produced the following results.

[RESULT] rows=2 ids=['A-001', 'A-002']
Ran 3 tests in 0.002s
OK

The three tests cover initial creation versus existing file replacement, retention of the original file when an error occurs mid-process, and reading CSV files where values contain newlines. Test processing times vary depending on the environment.

Write Process Sequence

  1. Create a temporary file in the same directory

  2. Write all records

  3. Detect column mismatches mid-process as errors

  4. Execute flush and fsync

  5. Close the temporary file

  6. Replace using os.replace

  7. Clean up the temporary file before replacement when an error occurs

Reason for creating in the same directory

Specify target.parent for the dir parameter of NamedTemporaryFile.

with tempfile.NamedTemporaryFile(
    mode="w",
    encoding="utf-8",
    newline="",
    delete=False,
    dir=target.parent
) as fp:
    ...

If a temporary file is created in the system default temp directory, it may reside on a different filesystem than the destination. Since os.replace can fail across different filesystems, it is created in the same directory.

Temporary files created with delete=False are not deleted simply when the program closes. Therefore, cleanup logic in a finally block is also required when an exception occurs. The specification can be found inPython tempfile.

Reason for setting newline to an empty string

In CSV files, cells may contain newlines. The Python csv module handles CSV newlines, such as by enclosing such values in quotes.

with csv_path.open(encoding="utf-8", newline="") as fp:
    rows = list(csv.DictReader(fp))

The official Pythoncsv documentationHowever, an example is shown where newline="" is specified when handling files. This prevents duplication between OS newline conversion and CSV-side processing.

Will the original CSV remain even if it is made to fail intentionally?

In the next test, an attempt is made to replace an existing CSV, but the second record has missing columns.

from pathlib import Path
from tempfile import TemporaryDirectory

with TemporaryDirectory() as folder:
    path = Path(folder) / "batch.csv"
    path.write_text("id,status\nOLD,OK\n", encoding="utf-8")

    try:
        replace_csv(path, (
            {"id": "A", "status": "DONE"},
            {"id": "MISSING"},
        ))
    except ValueError as error:
        print(f"[EXPECTED_ERROR] {error}")

    print(path.read_text(encoding="utf-8"))

The expected result is as follows.

[EXPECTED_ERROR] CSV row must have exactly id and status
id,status
OLD,OK

This is also verified in actual unit tests. Because an exception occurs before the replacement completes,the contents of the original CSV remain unchanged. The temporary file created during writing is deleted in the finally block.

However, this implementation does not continue processing by excluding only the failed records. The design aborts the replacement of the entire CSV if invalid input is detected.

Points to keep in mind in actual operation

Common situationApproach to handling
Another program also writes to itAdd file locking or similar mechanisms to control concurrency
The CSV is open in Excel on WindowsAnticipate the possibility of the replacement being denied and implement error handling and retry procedures
Disk space is insufficientStop as a write error and preserve the original data
Power failure occurs suddenlyConsider fsync, directory synchronization, and storage fault tolerance separately
The existing CSV contains important dataPerform backups and verify updates first
Handling large CSV filesEnsure sufficient disk space for temporary files

This methodis suited for use cases where a single writer safely updates a single CSV file.It does not guarantee database transactions, concurrency control, or recovery after a power outage.

When deploying to production, design the implementation to include directory permissions, file owners, existing file modes, backups, and rollbacks. Because using os.replace makes the new temporary file the replaced file, it is also important not to assume that the original file's access permissions and other attributes will be automatically inherited.

Referenced official documentation

Document information

Article title
Safely replacing CSV files in Python: Preserving original data even if write fails mid-process
Published
Updated
Source
https://papanda925.com/?p=18067&lang=en

License: Text and original figures for which this site holds the relevant rights are available under CC BY 4.0 , unless otherwise noted. This article may include content created or edited with generative AI. If code has a separate license notice or a linked GitHub repository license, that license takes precedence for the code. Quotations, third-party materials, images, and trademarks are excluded from this license. Usage policy

Copied title and URL