feat: add OpenSpec proposal for site content migration

Create comprehensive proposal to migrate all content from live markusgraf.ch
to the Hugo project, including text content, images, and structure.

- Add proposal.md with rationale and impact analysis
- Add design.md with content extraction strategy and trade-offs
- Add tasks.md with 13 ordered, verifiable implementation tasks
- Add content-migration spec with 5 requirements and 11 scenarios
- Add image-assets spec with 5 requirements and 10 scenarios

Validated with: openspec validate migrate-site-content --strict

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
Markus Graf
2025-10-27 10:58:48 +01:00
co-authored by Claude
parent c58c9ac9e6
commit 11401552b6
5 changed files with 334 additions and 0 deletions
@@ -0,0 +1,70 @@
# Design: Content Migration from markusgraf.ch
## Overview
This change migrates all existing content from the live markusgraf.ch website into the Hugo-based project structure. The migration is a one-time content extraction and conversion process.
## Approach
### Content Extraction Strategy
Since the live site cannot be automatically fetched, we rely on the user to provide:
1. HTML source code from the live site (via browser "Save As" or inspector)
2. Direct access to image URLs for downloading
### Content Conversion
- **HTML to Markdown**: Extract semantic content from HTML and convert to clean markdown
- **Preserve Structure**: Maintain the section hierarchy and flow from the live site
- **Hugo Integration**: Use Hugo's front matter and content organization patterns
### Image Handling
Images follow Hugo's static asset conventions:
- Store in `static/images/` (or appropriate subdirectory)
- Reference in markdown as `/images/filename.ext`
- Organize by purpose (e.g., `static/images/projects/`, `static/images/profile/`)
### Content Organization
```
content/
_index.md # Homepage with bio, intro, projects
about.md # About/CV page (if exists)
static/
images/
profile/ # Profile photos
projects/ # Project-related images
icons/ # Icons or small graphics (if any)
```
## Trade-offs
### Why Not Automate Extraction?
- WebFetch requires authentication (not available)
- Manual provision ensures accuracy and allows user to curate content
- One-time migration doesn't justify complex scraping infrastructure
### Content Structure Decisions
- **Single vs Multiple Pages**: If live site is single-page, we keep it as `_index.md`. If multi-page, we create separate content files
- **Project Organization**: Depending on number of projects, either:
- Inline in homepage (few projects)
- Separate section with list page (many projects)
### Image Optimization
- Accept images as-is from live site initially
- Future optimization (resizing, format conversion) can be separate change
- Prioritize content parity over optimization
## Dependencies
- Requires user to provide HTML source and ensure images are accessible
- Depends on existing Hugo site structure from `add-minimal-hugo-site` change
## Validation Strategy
- Visual comparison: Hugo build vs live site side-by-side
- Checklist verification: Every section, image, and link accounted for
- Browser console: No 404 errors or missing resources
- Accessibility: All images have alt text
## Rollback Plan
Since this is additive content migration:
- Git revert to restore placeholder content
- No breaking changes to existing Hugo structure
- Image assets can be removed from static/ directory
@@ -0,0 +1,15 @@
## Why
The Hugo site structure has been initialized but contains only placeholder content. To complete the migration from markusgraf.ch, we need to extract all actual content (text, images, and structure) from the live site and integrate it into the Hugo templates and content files. This ensures the Hugo site displays the exact same information as the current live site.
## What Changes
- Extract and migrate homepage content (bio/intro text) from markusgraf.ch
- Migrate all project/portfolio information and descriptions
- Download and organize all images from the live site into Hugo's static directory
- Update content markdown files (_index.md, etc.) with actual text content
- Ensure all internal links and image references work correctly in Hugo
- Verify visual and content parity between live site and Hugo build
## Impact
- Affected specs: content-migration (new), image-assets (new)
- Affected code: content/_index.md, content/about.md (if exists), static/ directory, potentially layout templates if content structure differs from placeholders
- Migration: One-time content extraction from live markusgraf.ch website
@@ -0,0 +1,82 @@
## ADDED Requirements
### Requirement: Content Extraction from Live Site
The system SHALL capture all text content, structure, and formatting from the live markusgraf.ch website to migrate into the Hugo project.
#### Scenario: Extract homepage content from live site
- **GIVEN** the live markusgraf.ch website HTML source is provided
- **WHEN** the HTML is parsed for content extraction
- **THEN** all text content, headings, and structure are captured
- **AND** the content is organized by logical sections (bio, intro, projects, etc.)
#### Scenario: Extract projects and portfolio content
- **GIVEN** the live site contains project or portfolio information
- **WHEN** projects are identified and extracted
- **THEN** each project has complete description captured
- **AND** project metadata and structure is documented
### Requirement: Content Conversion to Hugo Format
The system SHALL convert extracted HTML content into Hugo-compatible markdown format while preserving formatting and structure.
#### Scenario: Convert HTML to markdown
- **GIVEN** extracted HTML content from live site
- **WHEN** content is converted to markdown
- **THEN** all text formatting is preserved (bold, italic, links)
- **AND** HTML-specific elements are converted to markdown or Hugo shortcodes
- **AND** content follows markdown best practices
#### Scenario: Add Hugo front matter
- **GIVEN** converted markdown content
- **WHEN** content files are created
- **THEN** appropriate YAML front matter is added (title, description, date)
- **AND** front matter variables are correctly configured for templates
### Requirement: Homepage Content Integration
The system SHALL integrate migrated homepage content into content/_index.md, replacing placeholder content.
#### Scenario: Update homepage with actual content
- **GIVEN** converted homepage markdown content
- **WHEN** content/_index.md is updated
- **THEN** all placeholder text is replaced with actual content
- **AND** all sections from live site are present
- **AND** content hierarchy and flow matches live site structure
#### Scenario: Image references in content
- **GIVEN** homepage content references images
- **WHEN** image references are added to markdown
- **THEN** images use Hugo static path conventions (e.g., /images/photo.jpg)
- **AND** all image markdown syntax is correct
### Requirement: Content Structure Parity
The system SHALL ensure the Hugo site content structure matches the live markusgraf.ch site organization.
#### Scenario: Section organization matches live site
- **GIVEN** live site has distinct sections
- **WHEN** content is organized in Hugo
- **THEN** all sections are represented
- **AND** section order matches live site
- **AND** navigation between sections works correctly
#### Scenario: Multi-page structure if needed
- **GIVEN** live site has multiple pages
- **WHEN** pages are created in Hugo
- **THEN** each page has corresponding content file
- **AND** internal links between pages work
- **AND** navigation reflects page structure
### Requirement: Content Verification
The system SHALL verify that migrated content achieves parity with the live site.
#### Scenario: Content completeness check
- **GIVEN** Hugo site is built with migrated content
- **WHEN** compared with live markusgraf.ch
- **THEN** all text content from live site is present
- **AND** no content is missing or truncated
- **AND** content meaning and context are preserved
#### Scenario: Visual structure comparison
- **GIVEN** Hugo site is rendered
- **WHEN** viewed alongside live site
- **THEN** content sections appear in same order
- **AND** heading hierarchy matches
- **AND** overall content flow is equivalent
@@ -0,0 +1,86 @@
## ADDED Requirements
### Requirement: Image Asset Discovery
The system SHALL identify and catalog all images from the live markusgraf.ch website for migration.
#### Scenario: Identify all images on live site
- **GIVEN** the live markusgraf.ch website HTML source
- **WHEN** images are catalogued from the HTML
- **THEN** all image URLs are documented
- **AND** image purposes (profile, project, icon, etc.) are identified
- **AND** image file formats and sizes are noted
#### Scenario: Create asset inventory
- **GIVEN** identified images from live site
- **WHEN** asset inventory is created
- **THEN** a complete list of images with URLs exists
- **AND** each image is categorized by type/purpose
- **AND** inventory documents source URLs for downloading
### Requirement: Image Download and Organization
The system SHALL download images from the live site and organize them in Hugo's static directory following logical grouping principles.
#### Scenario: Download images from live site
- **GIVEN** list of image URLs from inventory
- **WHEN** images are downloaded
- **THEN** all images are successfully retrieved
- **AND** image files are verified for integrity
- **AND** no corrupted or failed downloads exist
#### Scenario: Organize images in static directory
- **GIVEN** downloaded images
- **WHEN** images are placed in Hugo project
- **THEN** images are saved to static/images/ or appropriate subdirectories
- **AND** directory structure reflects logical grouping (e.g., static/images/projects/, static/images/profile/)
- **AND** filenames are consistent and descriptive
### Requirement: Image Reference Integration
The system SHALL update all image references in content files to use Hugo's static path conventions correctly.
#### Scenario: Update image paths in content
- **GIVEN** images are in Hugo static directory
- **WHEN** content files reference images
- **THEN** image paths use Hugo static conventions (e.g., /images/photo.jpg)
- **AND** all image references use correct relative or absolute paths
- **AND** markdown image syntax is properly formatted
#### Scenario: Verify image links resolve
- **GIVEN** Hugo site is built
- **WHEN** pages with images are rendered
- **THEN** all image references resolve correctly
- **AND** no broken image links exist
- **AND** no 404 errors occur for image resources
### Requirement: Image Display Verification
The system SHALL ensure all migrated images display correctly in the built Hugo site.
#### Scenario: Images render correctly
- **GIVEN** Hugo site is built and served
- **WHEN** pages with images are viewed
- **THEN** all images display visually
- **AND** image aspect ratios are appropriate
- **AND** images load from correct static paths
- **AND** no missing or placeholder images appear
#### Scenario: Image dimensions and quality
- **GIVEN** images are displayed on site
- **WHEN** comparing to live site
- **THEN** image sizes are comparable to originals
- **AND** image quality is maintained
- **AND** no distortion or stretching occurs
### Requirement: Image Accessibility
The system SHALL ensure all images have appropriate accessibility attributes for screen readers and assistive technologies.
#### Scenario: Alt text for all images
- **GIVEN** images in content files
- **WHEN** markdown image syntax is used
- **THEN** all images include descriptive alt text
- **AND** alt text meaningfully describes image content
- **AND** decorative images use empty alt text where appropriate
#### Scenario: Semantic image usage
- **GIVEN** images serve specific purposes
- **WHEN** images are integrated into content
- **THEN** images are used semantically (figures, illustrations, etc.)
- **AND** image context is clear from surrounding content
@@ -0,0 +1,81 @@
# Tasks
## Content Extraction & Analysis
1. **Obtain HTML source from live markusgraf.ch**
- User provides saved HTML file or content extraction
- Parse and analyze the structure and sections
- Identify all text content blocks (bio, projects, etc.)
- **Validation**: HTML source file available and readable
2. **Catalog all images and assets**
- Extract all image URLs from live site HTML
- Document image filenames, paths, and purposes
- Create asset inventory list
- **Validation**: Complete list of images with URLs
## Image Migration
3. **Download all images from live site**
- Download each catalogued image
- Verify image integrity (no corrupted downloads)
- **Validation**: All images downloaded successfully, checksums verify
4. **Organize images in static directory**
- Create appropriate subdirectories in static/ (e.g., static/images/)
- Place images with consistent naming
- Document the organization structure
- **Validation**: Images accessible via Hugo static path convention
## Content Migration
5. **Extract and convert homepage content**
- Parse homepage HTML to identify content sections
- Convert HTML to markdown format
- Preserve formatting (bold, italic, links, etc.)
- **Validation**: Markdown content matches original HTML structure
6. **Update _index.md with actual content**
- Replace placeholder content in content/_index.md
- Add proper front matter
- Integrate all homepage sections
- Update image references to Hugo paths
- **Validation**: `hugo server` displays homepage content correctly
7. **Migrate projects/portfolio content**
- Extract project descriptions and details
- Create appropriate content structure (individual files or sections)
- Add project images and link them correctly
- **Validation**: All projects visible with complete information
8. **Update or create additional content pages**
- Create/update about page if present on live site
- Ensure all pages from live site are represented
- **Validation**: Page structure matches live site
## Verification & Quality Assurance
9. **Verify all internal links**
- Check all links between pages work
- Verify anchor links if used
- **Validation**: No broken internal links, `hugo server` navigation works
10. **Verify all image references**
- Build Hugo site
- Check that all images display correctly
- Verify no 404 errors in browser console
- **Validation**: All images load, no missing assets
11. **Content parity verification**
- Compare Hugo build output with live markusgraf.ch
- Verify all text content is present
- Check that layout and sections match
- **Validation**: Side-by-side comparison confirms identical content
12. **Accessibility check for images**
- Add alt text to all images
- Ensure image references are semantic
- **Validation**: All `<img>` tags have descriptive alt attributes
## Documentation
13. **Document content structure**
- Update project documentation with content organization
- Note any deviations from live site (if applicable)
- Document image organization scheme
- **Validation**: README or docs reflect current content structure