Complete Phase 0 (Research) and Phase 1 (Design) of implementation planning workflow: - plan.md: Technical context, constitution check, and project structure - research.md: Technology decisions for AI integration, file storage, authentication, security - data-model.md: Domain entities with validation rules and state transitions - contracts/: API specifications for submission, dashboard, and admin endpoints - quickstart.md: Developer setup guide with test-first workflow - CLAUDE.md: Updated agent context with tech stack All constitutional requirements validated. Ready for task generation phase. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
14 KiB
Research: Anonymous Feedback Platform (Reklamator)
Branch: 001-build-an-application | Date: 2025-10-15
This document resolves all NEEDS CLARIFICATION items identified in the Technical Context section of plan.md.
1. AI Integration Approach
Decision: Pluggable AI provider interface with Claude as default
Rationale:
- User input specifies "we use claude but it can be any other service as well"
- Designing for extensibility aligns with good architectural practice
- Enables future migration to different AI providers without major refactoring
Implementation Approach:
- Abstract base class
AIAnalyzerdefining interface:analyze_feedback(text: str, target_lang: str) -> AnalysisResult - Concrete implementation
ClaudeAnalyzerusing Anthropic API - Configuration-driven provider selection
- API key management via environment variables
Alternatives Considered:
- Hard-coded Claude API integration: Simpler initially but violates user requirement for provider flexibility
- LangChain framework: Adds significant dependency weight for simple translation/categorization task
- Multiple provider implementations from start: Premature complexity - implement Claude first, abstract as needed
Claude API Specifics:
- Use
anthropicPython SDK - Model:
claude-3-haiku-20240307for cost-effective analysis (fast, sufficient for categorization/translation) - Prompt design: Single API call with structured output for category, summary, translation
- Error handling: Retry logic for transient failures, graceful degradation for persistent errors
Research References:
- Anthropic API Documentation: https://docs.anthropic.com/
- Python SDK: https://github.com/anthropics/anthropic-sdk-python
2. File Upload - Malware Scanning Approach
Decision: ClamAV integration via clamd for virus scanning
Rationale:
- FR-060 requires malware scanning before storage
- ClamAV is open-source, widely used, actively maintained
clamdprovides Python bindings for integration- Suitable for on-premise deployment matching file-based storage philosophy
Implementation Approach:
- Install ClamAV daemon (
clamd) as system service - Use
clamdPython library for scanning uploaded files - Scan files synchronously during upload before writing to disk
- Reject files that fail virus scan with clear error message
- Log scanning failures for security monitoring
Configuration:
- Maximum file size: 10MB per file (FR-006)
- Allowed extensions:
.pdf,.docx,.txt,.jpg,.png,.gif,.webp - MIME type validation in addition to extension checking
- Temporary upload storage cleaned after scan (pass or fail)
Alternatives Considered:
- Cloud-based scanning (VirusTotal API): Violates anonymity requirement (uploads data externally), adds latency
- No scanning: Violates FR-060 security requirement
- Manual review: Not scalable, delays feedback processing
- Python-based scanning (yara-python): More complex to configure, less comprehensive than ClamAV
Dependencies:
clamdPython library- ClamAV daemon installed on server
3. File Upload - Storage Location Strategy
Decision: Local filesystem storage in data/products/{product-id}/feedback/{feedback-id}/attachments/
Rationale:
- Aligns with file-based storage architecture (no database)
- User input specifies "folders foreach user input" and "files in a folder"
- Simple to implement, backup, and inspect
- No additional service dependencies
- Sufficient for MVP scale (100 products × 10k items × 3 files × 10MB = ~30TB worst case)
Directory Structure:
data/
└── products/
└── {product-id}/ # e.g., "001-acme-app"
├── config.yaml # Product metadata (name, target language, owners)
└── feedback/
└── {feedback-id}/ # UUID v4, e.g., "a3f2c1d5-..."
├── metadata.yaml # Feedback metadata (timestamp, status, category, etc.)
├── content.txt # Original feedback text
├── analysis.md # AI-generated analysis report
└── attachments/
├── original_filename_1.pdf
├── original_filename_2.png
└── original_filename_3.jpg
File Naming:
- Preserve original filenames to maintain user context
- Sanitize filenames to prevent directory traversal (strip
../, absolute paths, etc.) - Handle duplicate filenames by appending counter if needed
Alternatives Considered:
- Cloud storage (S3/GCS): Adds external dependency, cost, complexity; overkill for MVP
- Flat directory per product: Poor scalability, difficult to organize metadata
- Database with BLOB storage: Contradicts file-based storage decision, adds DB complexity
- Content-addressed storage (hash-based filenames): Loses original filename context, complicates presentation
Backup Strategy (out of scope for MVP but noted):
- Simple filesystem backup via rsync/tar sufficient
- Can upgrade to cloud sync if needed later
4. Authentication - Session Management Approach
Decision: Flask-Login with server-side sessions for product owners/admins
Rationale:
- Flask-Login is standard, well-tested session management for Flask
- Server-side sessions prevent token tampering
- Simple username/password authentication sufficient for MVP
- FR-056, FR-063 require authentication and secure password storage
Implementation Approach:
- Use
Flask-Loginextension for session management - Store user credentials in simple YAML file (products/users.yaml) for MVP consistency with file-based approach
- Hash passwords with
bcrypt(FR-063) - Session cookies:
HttpOnly,Secure(HTTPS only),SameSite=Lax - Session timeout: 24 hours of inactivity
User Model:
users:
- id: "admin-001"
email: "admin@example.com"
password_hash: "$2b$12$..."
role: "admin"
assigned_products: [] # Empty = all products access
- id: "owner-001"
email: "owner@example.com"
password_hash: "$2b$12$..."
role: "product_owner"
assigned_products: ["001-acme-app", "002-beta-service"]
Access Control:
- Admins: Full access to all products, can manage products/owners
- Product Owners: Read-only access to assigned products only (FR-033)
- Anonymous users: Submission form access only (no authentication)
Alternatives Considered:
- JWT tokens: More complex, unnecessary for server-rendered HTML application
- OAuth/SAML: Over-engineered for MVP, adds external identity provider dependency
- Database-backed sessions: Contradicts file-based architecture
- No authentication: Violates FR-056 requirement
Dependencies:
Flask-Loginextensionbcryptfor password hashing
5. Rate Limiting - Implementation Strategy
Decision: Flask-Limiter with IP-based rate limiting for submission endpoint
Rationale:
- FR-061 requires rate limiting (suggested: 10 submissions/hour/IP)
- Flask-Limiter is standard, well-maintained Flask extension
- IP-based limiting suitable for anonymous submissions
- In-memory storage sufficient for MVP (single server deployment)
Implementation Approach:
- Use
Flask-Limiterextension - Apply rate limit decorator to submission route:
@limiter.limit("10 per hour") - Storage backend: In-memory (default) for MVP
- Return HTTP 429 Too Many Requests with clear error message
- Exempt authenticated admin users from rate limits (for testing)
Rate Limit Configuration:
- Anonymous submission: 10 requests per hour per IP address
- Dashboard/admin routes: No rate limiting (authenticated users only)
- Rate limit headers included in response:
X-RateLimit-Limit,X-RateLimit-Remaining,X-RateLimit-Reset
Considerations:
- IP-based limiting can be circumvented via VPN/proxy but sufficient deterrent for casual abuse
- Behind proxy/load balancer: Configure Flask-Limiter to read
X-Forwarded-Forheader - Note: FR-055 requires no fingerprinting for identification - rate limiting is for abuse prevention only, not user tracking
Alternatives Considered:
- CAPTCHA (hCaptcha/reCAPTCHA): Adds friction to user experience, contradicts "lower barriers" goal
- Redis-backed rate limiting: Unnecessary complexity for single-server MVP
- No rate limiting: Violates FR-061 requirement, leaves system vulnerable to abuse
- Token bucket per session: Requires session tracking for anonymous users, violates anonymity
Dependencies:
Flask-Limiterextension
6. Flask Best Practices for Simple HTML Applications
Decision: Server-side rendering with Jinja2 templates, no JavaScript
Rationale:
- User explicitly specifies "plain html using flask" and "does not use any css frameworks or javascript libraries"
- Server-side rendering eliminates frontend build complexity
- Jinja2 included with Flask, no additional dependencies
- Forms use standard HTTP POST/GET, progressive enhancement approach
Template Approach:
- Minimal inline CSS for basic layout (no framework)
- Semantic HTML5 for accessibility
- Server-side form validation with error display
- Standard browser form controls (no custom widgets)
Form Handling:
- POST requests for submissions
- Server-side validation with error messages
- Flash messages for user feedback
- Redirect-after-POST pattern to prevent duplicate submissions
No JavaScript Requirement (Edge Case from spec.md line 105):
- "What happens when a user's browser doesn't support JavaScript"
- Answer: Application works fully without JavaScript (no JS used)
- File uploads work via standard HTML
<input type="file" multiple>
Best Practices Applied:
- Flask app factory pattern for testability
- Blueprint organization for routes
- Environment-based configuration
- CSRF protection via Flask-WTF (even for simple forms)
Dependencies:
- Flask (includes Jinja2)
- Flask-WTF for CSRF protection
7. Python Dependency Management
Decision: requirements.txt with pinned versions for reproducibility
Rationale:
- Simplest dependency management for Flask application
- No need for Poetry/Pipenv complexity in MVP
- Pin exact versions for reproducibility
- Virtual environment assumed for isolation
Core Dependencies (estimated):
Flask==3.0.0
Flask-Login==0.6.3
Flask-Limiter==3.5.0
Flask-WTF==1.2.1
anthropic==0.8.0
clamd==1.0.2
bcrypt==4.1.2
PyYAML==6.0.1
pytest==7.4.3
pytest-flask==1.3.0
Development Dependencies:
- pytest, pytest-flask for testing
- black for code formatting
- flake8 for linting
Technology Stack Summary
| Component | Technology | Rationale |
|---|---|---|
| Web Framework | Flask 3.0+ | Lightweight, simple, widely supported |
| Template Engine | Jinja2 (built-in) | Server-side rendering, no JS needed |
| AI Provider | Claude API (Anthropic) | User-specified, abstracted for future flexibility |
| Authentication | Flask-Login + bcrypt | Standard session management, secure passwords |
| Rate Limiting | Flask-Limiter | Prevent abuse, simple IP-based approach |
| Malware Scanning | ClamAV + clamd | Open-source, reliable, on-premise |
| Storage | Filesystem (YAML + Markdown) | Matches user requirements, simple, no DB |
| Testing | pytest + pytest-flask | Industry standard, good Flask integration |
| Python Version | 3.11+ | Modern, stable, good performance |
| Deployment | Gunicorn (WSGI) | Production-ready Flask server |
Non-Functional Requirements Research
Performance Considerations
Concurrent Submissions (SC-012):
- Target: 100 concurrent submissions without errors
- Flask + Gunicorn with 4-8 worker processes should handle this
- File I/O is bottleneck: consider async I/O if performance issues arise
- AI analysis happens asynchronously (background task) to not block submission response
Dashboard Performance (SC-008):
- Target: Load 1000 items in <3 seconds
- File-based approach: Index product feedback directories, cache counts
- Implement pagination (50 items per page)
- Use lazy loading for file attachments (links, not embedded content)
AI Analysis Time (SC-007):
- Target: <30 seconds for 95% of submissions
- Claude Haiku model typically responds in 2-5 seconds for translation/categorization
- Timeout: 45 seconds before marking as failed
- Queue-based processing if needed (Python
queuemodule or simple file-based queue)
Security Considerations
Anonymity Enforcement (FR-055, SC-010):
- Do NOT log IP addresses in feedback metadata
- Rate limiting uses IP for abuse prevention only, not stored with feedback
- No session cookies for anonymous submission
- No analytics/tracking scripts
File Upload Security:
- Validate MIME types server-side (don't trust client)
- ClamAV scanning before storage
- Sanitize filenames to prevent directory traversal
- Store outside web root, serve via Flask route with access control
HTTPS Requirement (FR-064):
- Deployment guide must specify reverse proxy (nginx) with TLS
- Redirect HTTP to HTTPS
- HSTS headers recommended
Open Questions for Implementation Phase
-
Asynchronous AI Analysis: Should analysis happen synchronously (user waits) or asynchronously (background job)?
- Recommendation: Asynchronous - return success immediately, process in background
- Implement simple file-based queue or use Python
threadingfor MVP
-
Admin Bootstrap: How is the first admin user created?
- Recommendation: CLI command or config file initialization script
-
Email Notifications: Out of scope (line 255) but commonly requested
- Recommendation: Document as future enhancement, design hooks for extensibility
-
Logging Strategy: Structured logs for operational monitoring?
- Recommendation: Python
loggingmodule, JSON format, separate file per environment
- Recommendation: Python
Research Validation
All NEEDS CLARIFICATION items from Technical Context have been resolved:
| Item | Resolution | Document Section |
|---|---|---|
| AI Integration | Pluggable interface, Claude as default | §1 |
| File Upload - Malware Scanning | ClamAV + clamd | §2 |
| File Upload - Storage Location | Filesystem: data/products/.../feedback/.../attachments/ |
§3 |
| Authentication | Flask-Login + server-side sessions + bcrypt | §4 |
| Rate Limiting | Flask-Limiter, 10/hour/IP | §5 |
Next Phase: Proceed to Phase 1 (data-model.md, contracts, quickstart.md)