| Description | With $g_file_upload_method = DISK, every attachment belonging to a project is
written directly into that project's upload directory. There is no way to
subdivide it.
On installations that accumulate large numbers of attachments this becomes a
real operational problem. In our case a single directory holds roughly 133,000
files. The consequences:
- Directory traversal is expensive, and every tool that walks the directory
pays for it.
- Incremental backups are the worst case. rsync has to stat the entire
directory to discover the handful of files that changed since the last run,
which took hours. After splitting the files across 256 subdirectories the
same backup completes in minutes, because unchanged subtrees are skipped
cheaply.
Other applications that store large numbers of hashed files solve this by
using the leading characters of the file name as a directory prefix - git's
object store being the obvious example. |
|---|
| Steps To Reproduce | Two new configuration options:
$g_file_upload_subdirectory_depth, defaulting to 0, which preserves the
current flat layout exactly. A depth of 1 gives 256 subdirectories, a depth
of 2 gives 65536.
$g_file_upload_subdirectory_width, defaulting to 2, the number of
characters consumed per level.
Deriving the subdirectory from the disk file name itself, rather than from a
separate hash, keeps the mapping reversible: a file can still be located from
its name alone without consulting the database, which matters when restoring
from backup.
No schema change or data migration is required. The folder column already
records each attachment's location per row, so files written under the old and
new layouts coexist without any special handling. (This does depend on the
readers actually honouring that column - reported separately.)
An admin script can relocate attachments already on disk, in either direction,
for installations that want to adopt the new layout for existing data. |
|---|
| Additional Information | One caveat worth recording: disk file names are normally 32 character hashes,
but installations that have imported data from another tracker can hold other
formats. Ours contains UUID-style names, which offer only 8 hexadecimal
characters before the first separator. Any implementation needs to use only
the leading hexadecimal run and fall back to the flat layout for names too
short to satisfy the configured depth, rather than assuming a 32 character
hash. |
|---|