Building Gemini Spark Skills
A Gemini Spark Skill is a JSON-defined unit of work with: a schema (inputs/outputs), reusable instructions (the "prompt" logic with context), optional app connections, and a trigger (manual, scheduled, or event-based).
Minimal skill definition:
JSON{ "name": "document_summarizer", "description": "Summarizes PDF, DOCX, or PPTX files into a concise Markdown brief", "input_schema": { "type": "object", "properties": { "file": { "type": "string", "format": "file", "accepts": ["pdf", "docx", "pptx"] }, "max_words": { "type": "integer", "default": 300 } }, "required": ["file"] }, "output_schema": { "type": "object", "properties": { "summary_markdown": { "type": "string" } } }, "instructions": "Extract text and structure from {{file}}. Preserve headings and key figures. Produce a Markdown summary under {{max_words}} words. Flag any tables or images that could not be parsed.", "connections": [], "trigger": { "type": "manual" } }
Test it locally with the Spark SDK before wiring connections or schedules:
Bashspark run --skill document_summarizer.json --input sample.pdf
Progress:
- Step 1: Define the JSON schema (inputs, outputs, types, defaults, required fields)
- Step 2: Write instructions — the reusable prompt logic, using
{{variable}}interpolation and explicit edge-case handling - Step 3: Add context — examples, format rules, failure fallbacks, style constraints
- Step 4: Connect Workspace/apps (Drive, Gmail, Sheets, GitHub, Slack) only if the skill needs external I/O
- Step 5: Test locally with representative + adversarial inputs (malformed files, empty input, oversized files)
- Step 6: Set a trigger — manual, time-based (cron), or condition-based (file arrival, webhook)
- Step 7: Deploy and monitor — check logs, error rates, and output quality on a sample of live runs
- Step 8: Chain into pipelines if the skill is one stage of a multi-skill workflow
- JSON Schema — strict typing for inputs/outputs. Always set
required, sensibledefaults, andformathints (file,date,enum) so Spark validates before invoking the model. - Instructions — the reusable "program" written in natural language + interpolation. Treat this like a function body: define inputs, transformation steps, and the exact output contract.
- Context — few-shot examples, domain vocabulary, style/format rules, and explicit "if X happens, do Y" fallback rules. This is what separates a fragile skill from a robust one.
- Trigger + Connections — how the skill starts (manual/schedule/event) and what external systems it touches (Drive folder watch, Gmail label, GitHub repo webhook, Sheet row append).
Example 1: Format conversion skill Input: A skill spec request — "convert any Office/CSV file to clean Markdown" Output:
JSON{ "name": "any_doc_to_markdown", "input_schema": { "type": "object", "properties": { "file": { "type": "string", "format": "file", "accepts": ["docx","xlsx","pptx","csv","odt"] } }, "required": ["file"] }, "output_schema": { "type": "object", "properties": { "markdown": { "type": "string" } } }, "instructions": "Detect file type from {{file}} extension. Extract text preserving heading hierarchy and tables as GFM tables. For CSV, render as a single Markdown table with the first row as header. If the file has embedded images, insert a placeholder `[image: description]` instead of failing. Output only the Markdown body, no commentary.", "trigger": { "type": "manual" } }
Example 2: Autonomous scheduled skill Input: "Transcribe meeting audio dropped into a Drive folder every night" Output: Add a Drive-watch connection plus a cron trigger:
JSON{ "connections": [{ "app": "google_drive", "watch_folder": "MeetingAudio/", "on": "file_added" }], "trigger": { "type": "schedule", "cron": "0 2 * * *", "condition": "new_files_since_last_run" } }
Instructions should explicitly state output destination (e.g., write .md transcript back to a Transcripts/ folder) so the skill is fully autonomous with no manual pickup step.
- One skill, one responsibility. Don't build a skill that both converts and emails and summarizes — chain three skills instead. Easier to debug and reuse.
- Always define output contracts precisely. Ambiguous instructions ("summarize the file") produce inconsistent output; specify format, length, and structure.
- Add explicit fallback instructions for malformed/partial input ("if OCR confidence is low, mark the line with
[unclear]rather than guessing"). - Version your schemas. When you change input/output shape, bump a
versionfield so chained skills don't silently break. - Test with adversarial inputs before scheduling. Empty files, huge files, wrong formats — a skill running 24/7 unattended will eventually see all of these.
- Least-privilege connections. Only grant the Workspace/app scopes the skill actually needs (e.g., read-only Drive access for a converter that never writes back).
- Log inputs/outputs for scheduled skills. You can't interactively debug something that ran at 2 a.m. on a device that was off — logs are the only visibility you'll have.
- Chain via well-defined intermediate formats (Markdown, JSON, Parquet) — this is what makes 30 individual skills composable into one pipeline.
- Vague instructions with no output contract — leads to non-deterministic formatting that breaks downstream skills in a chain.
- Skipping local testing before scheduling — a skill with a subtle bug run every night at 2 a.m. can silently corrupt/lose data for days before anyone notices.
- Over-broad app permissions — requesting full Gmail/Drive access when the skill only needs to read one folder is a security liability.
- No error handling for unsupported formats — a converter that throws an unhandled exception on a corrupt PDF halts the whole autonomous pipeline instead of flagging and continuing.
- Hardcoding values that belong in the schema — e.g., baking a folder path or max-word count into instructions instead of exposing it as a schema parameter, which kills reusability.
- Chaining skills without a shared intermediate contract — if skill A outputs prose and skill B expects strict JSON, the pipeline breaks silently.
- Forgetting to monitor after deployment — "set and forget" scheduled skills need periodic log review; silent failures compound over time in 24/7 operation.