Curating Research to Markdown
Pythondef research_curator(item): if not authorized(item): return {"status": "skipped", "reason": "Access not authorized"} kind = detect_format(item) if kind not in SUPPORTED_INPUTS: return {"status": "skipped", "reason": "Unsupported format"} result = markitdown_convert(item) output = add_metadata( result.markdown, source=item.url, processed_at=current_timestamp(), warnings=result.warnings, ) save_markdown(item.parent_folder, item.stem + ".md", output) return {"status": "converted", "warnings": result.warnings}
Run this against one file per supported type before enabling unattended/batch processing.
Progress:
- Confirm MarkItDown's current supported formats and dependencies (formats/behavior change between versions)
- Identify watched folder or explicit submission channel (Drive folder, connected app URL)
- Detect input format by content inspection, not filename extension alone
- Check authorization/access control before touching any file or link
- Convert eligible files via MarkItDown, preserving headings, tables, links, timestamps
- Handle ZIP archives: list contents, convert eligible members individually, never execute anything inside
- Handle YouTube URLs: pull captions/transcripts only via authorized methods; record when unavailable
- Write output as
<original-stem>.mdbeside the source, with metadata block (source, processed_at, detected format, warnings) - Report status explicitly: converted / partial / skipped / failed — never silently assume success
- Leave original source file untouched
Example 1:
Input: quarterly-report.pdf dropped in watched Drive folder, containing headings, a data table, and one embedded image.
Output: quarterly-report.md saved in the same folder, with headings and table preserved as Markdown, a note that the image was omitted (warnings: ["image content not extracted"]), plus a metadata footer listing source path and processing timestamp.
Example 2:
Input: A YouTube URL submitted through the connected app, video has no captions available.
Output: <video-title>.md containing only title/description metadata and an explicit warnings: ["no captions or transcript available"] line — not a fabricated transcript.
Example 3:
Input: assets.zip containing notes.docx, data.xlsx, and installer.exe.
Output: assets.zip.md listing all three archive contents; notes.md and data.md produced from the eligible files; installer.exe explicitly flagged as skipped ("unsupported/executable, not run").
- Treat conversion as extraction, not verification — Markdown output may lose layout, images, formulas, embedded objects. State this limitation in output metadata when relevant.
- Detect format from actual content/magic bytes when possible; a
.docxextension on a corrupted or renamed file should not be trusted blindly. - Always preserve a link/reference back to the original source file or URL in the output.
- Keep a per-item status record (converted/partial/skipped/failed) rather than a single pipeline-level success flag.
- Apply the same access-control checks to output folders as to source folders — converted Markdown can leak confidential content just as easily as the original.
- For batch/ZIP processing, convert each eligible member independently so one bad file doesn't abort the whole batch.
- Silent failure: never report "converted" when MarkItDown returned partial content or an error — surface warnings explicitly.
- Executing archive contents: ZIP members must only be listed/converted, never run, even if they look like documents.
- Inferring format from filename: a mislabeled or renamed file can crash or mis-convert; verify content type first.
- Fabricating missing content: if captions, transcripts, formulas, or images aren't extractable, say so — don't approximate or invent substitute text.
- Overwriting originals: output must be a sibling
.mdfile, never replace or modify the source. - Processing without authorization checks: confirm access rights before converting anything, especially for shared drives or third-party URLs.