Creative archives accumulate sideways. A project folder becomes a hard drive, then three hard drives, then a cloud bucket full of “final,” “final-final,” and files whose only metadata is the memory of the person who made them.
The pipeline
A useful first version can stay deliberately boring: walk the directories, normalize filenames without destroying originals, inspect technical metadata, calculate hashes, identify probable duplicates, classify assets, and write a searchable inventory.
What the inventory should know
- Original path and stable asset identifier
- Filename history and normalized display name
- File type, dimensions, duration, codec, and size
- Creation and modification dates
- Duplicate or near-duplicate relationships
- Project, campaign, or collection membership
- Human-confirmed tags versus AI-generated suggestions
The intelligence layer becomes useful only after the evidence layer is reliable. If an assistant cannot show which file a tag came from, it is not asset intelligence yet—it is decoration.
Start with the inventory. Add the agent after the archive can tell the truth.