Add implementation planning artifacts for anonymous feedback platform
Complete Phase 0 (Research) and Phase 1 (Design) of implementation planning workflow: - plan.md: Technical context, constitution check, and project structure - research.md: Technology decisions for AI integration, file storage, authentication, security - data-model.md: Domain entities with validation rules and state transitions - contracts/: API specifications for submission, dashboard, and admin endpoints - quickstart.md: Developer setup guide with test-first workflow - CLAUDE.md: Updated agent context with tech stack All constitutional requirements validated. Ready for task generation phase. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,366 @@
|
||||
# Research: Anonymous Feedback Platform (Reklamator)
|
||||
|
||||
**Branch**: `001-build-an-application` | **Date**: 2025-10-15
|
||||
|
||||
This document resolves all NEEDS CLARIFICATION items identified in the Technical Context section of plan.md.
|
||||
|
||||
## 1. AI Integration Approach
|
||||
|
||||
### Decision: Pluggable AI provider interface with Claude as default
|
||||
|
||||
**Rationale**:
|
||||
- User input specifies "we use claude but it can be any other service as well"
|
||||
- Designing for extensibility aligns with good architectural practice
|
||||
- Enables future migration to different AI providers without major refactoring
|
||||
|
||||
**Implementation Approach**:
|
||||
- Abstract base class `AIAnalyzer` defining interface: `analyze_feedback(text: str, target_lang: str) -> AnalysisResult`
|
||||
- Concrete implementation `ClaudeAnalyzer` using Anthropic API
|
||||
- Configuration-driven provider selection
|
||||
- API key management via environment variables
|
||||
|
||||
**Alternatives Considered**:
|
||||
- **Hard-coded Claude API integration**: Simpler initially but violates user requirement for provider flexibility
|
||||
- **LangChain framework**: Adds significant dependency weight for simple translation/categorization task
|
||||
- **Multiple provider implementations from start**: Premature complexity - implement Claude first, abstract as needed
|
||||
|
||||
**Claude API Specifics**:
|
||||
- Use `anthropic` Python SDK
|
||||
- Model: `claude-3-haiku-20240307` for cost-effective analysis (fast, sufficient for categorization/translation)
|
||||
- Prompt design: Single API call with structured output for category, summary, translation
|
||||
- Error handling: Retry logic for transient failures, graceful degradation for persistent errors
|
||||
|
||||
**Research References**:
|
||||
- Anthropic API Documentation: https://docs.anthropic.com/
|
||||
- Python SDK: https://github.com/anthropics/anthropic-sdk-python
|
||||
|
||||
---
|
||||
|
||||
## 2. File Upload - Malware Scanning Approach
|
||||
|
||||
### Decision: ClamAV integration via clamd for virus scanning
|
||||
|
||||
**Rationale**:
|
||||
- FR-060 requires malware scanning before storage
|
||||
- ClamAV is open-source, widely used, actively maintained
|
||||
- `clamd` provides Python bindings for integration
|
||||
- Suitable for on-premise deployment matching file-based storage philosophy
|
||||
|
||||
**Implementation Approach**:
|
||||
- Install ClamAV daemon (`clamd`) as system service
|
||||
- Use `clamd` Python library for scanning uploaded files
|
||||
- Scan files synchronously during upload before writing to disk
|
||||
- Reject files that fail virus scan with clear error message
|
||||
- Log scanning failures for security monitoring
|
||||
|
||||
**Configuration**:
|
||||
- Maximum file size: 10MB per file (FR-006)
|
||||
- Allowed extensions: `.pdf`, `.docx`, `.txt`, `.jpg`, `.png`, `.gif`, `.webp`
|
||||
- MIME type validation in addition to extension checking
|
||||
- Temporary upload storage cleaned after scan (pass or fail)
|
||||
|
||||
**Alternatives Considered**:
|
||||
- **Cloud-based scanning (VirusTotal API)**: Violates anonymity requirement (uploads data externally), adds latency
|
||||
- **No scanning**: Violates FR-060 security requirement
|
||||
- **Manual review**: Not scalable, delays feedback processing
|
||||
- **Python-based scanning (yara-python)**: More complex to configure, less comprehensive than ClamAV
|
||||
|
||||
**Dependencies**:
|
||||
- `clamd` Python library
|
||||
- ClamAV daemon installed on server
|
||||
|
||||
---
|
||||
|
||||
## 3. File Upload - Storage Location Strategy
|
||||
|
||||
### Decision: Local filesystem storage in `data/products/{product-id}/feedback/{feedback-id}/attachments/`
|
||||
|
||||
**Rationale**:
|
||||
- Aligns with file-based storage architecture (no database)
|
||||
- User input specifies "folders foreach user input" and "files in a folder"
|
||||
- Simple to implement, backup, and inspect
|
||||
- No additional service dependencies
|
||||
- Sufficient for MVP scale (100 products × 10k items × 3 files × 10MB = ~30TB worst case)
|
||||
|
||||
**Directory Structure**:
|
||||
```
|
||||
data/
|
||||
└── products/
|
||||
└── {product-id}/ # e.g., "001-acme-app"
|
||||
├── config.yaml # Product metadata (name, target language, owners)
|
||||
└── feedback/
|
||||
└── {feedback-id}/ # UUID v4, e.g., "a3f2c1d5-..."
|
||||
├── metadata.yaml # Feedback metadata (timestamp, status, category, etc.)
|
||||
├── content.txt # Original feedback text
|
||||
├── analysis.md # AI-generated analysis report
|
||||
└── attachments/
|
||||
├── original_filename_1.pdf
|
||||
├── original_filename_2.png
|
||||
└── original_filename_3.jpg
|
||||
```
|
||||
|
||||
**File Naming**:
|
||||
- Preserve original filenames to maintain user context
|
||||
- Sanitize filenames to prevent directory traversal (strip `../`, absolute paths, etc.)
|
||||
- Handle duplicate filenames by appending counter if needed
|
||||
|
||||
**Alternatives Considered**:
|
||||
- **Cloud storage (S3/GCS)**: Adds external dependency, cost, complexity; overkill for MVP
|
||||
- **Flat directory per product**: Poor scalability, difficult to organize metadata
|
||||
- **Database with BLOB storage**: Contradicts file-based storage decision, adds DB complexity
|
||||
- **Content-addressed storage (hash-based filenames)**: Loses original filename context, complicates presentation
|
||||
|
||||
**Backup Strategy** (out of scope for MVP but noted):
|
||||
- Simple filesystem backup via rsync/tar sufficient
|
||||
- Can upgrade to cloud sync if needed later
|
||||
|
||||
---
|
||||
|
||||
## 4. Authentication - Session Management Approach
|
||||
|
||||
### Decision: Flask-Login with server-side sessions for product owners/admins
|
||||
|
||||
**Rationale**:
|
||||
- Flask-Login is standard, well-tested session management for Flask
|
||||
- Server-side sessions prevent token tampering
|
||||
- Simple username/password authentication sufficient for MVP
|
||||
- FR-056, FR-063 require authentication and secure password storage
|
||||
|
||||
**Implementation Approach**:
|
||||
- Use `Flask-Login` extension for session management
|
||||
- Store user credentials in simple YAML file (products/users.yaml) for MVP consistency with file-based approach
|
||||
- Hash passwords with `bcrypt` (FR-063)
|
||||
- Session cookies: `HttpOnly`, `Secure` (HTTPS only), `SameSite=Lax`
|
||||
- Session timeout: 24 hours of inactivity
|
||||
|
||||
**User Model**:
|
||||
```yaml
|
||||
users:
|
||||
- id: "admin-001"
|
||||
email: "admin@example.com"
|
||||
password_hash: "$2b$12$..."
|
||||
role: "admin"
|
||||
assigned_products: [] # Empty = all products access
|
||||
|
||||
- id: "owner-001"
|
||||
email: "owner@example.com"
|
||||
password_hash: "$2b$12$..."
|
||||
role: "product_owner"
|
||||
assigned_products: ["001-acme-app", "002-beta-service"]
|
||||
```
|
||||
|
||||
**Access Control**:
|
||||
- Admins: Full access to all products, can manage products/owners
|
||||
- Product Owners: Read-only access to assigned products only (FR-033)
|
||||
- Anonymous users: Submission form access only (no authentication)
|
||||
|
||||
**Alternatives Considered**:
|
||||
- **JWT tokens**: More complex, unnecessary for server-rendered HTML application
|
||||
- **OAuth/SAML**: Over-engineered for MVP, adds external identity provider dependency
|
||||
- **Database-backed sessions**: Contradicts file-based architecture
|
||||
- **No authentication**: Violates FR-056 requirement
|
||||
|
||||
**Dependencies**:
|
||||
- `Flask-Login` extension
|
||||
- `bcrypt` for password hashing
|
||||
|
||||
---
|
||||
|
||||
## 5. Rate Limiting - Implementation Strategy
|
||||
|
||||
### Decision: Flask-Limiter with IP-based rate limiting for submission endpoint
|
||||
|
||||
**Rationale**:
|
||||
- FR-061 requires rate limiting (suggested: 10 submissions/hour/IP)
|
||||
- Flask-Limiter is standard, well-maintained Flask extension
|
||||
- IP-based limiting suitable for anonymous submissions
|
||||
- In-memory storage sufficient for MVP (single server deployment)
|
||||
|
||||
**Implementation Approach**:
|
||||
- Use `Flask-Limiter` extension
|
||||
- Apply rate limit decorator to submission route: `@limiter.limit("10 per hour")`
|
||||
- Storage backend: In-memory (default) for MVP
|
||||
- Return HTTP 429 Too Many Requests with clear error message
|
||||
- Exempt authenticated admin users from rate limits (for testing)
|
||||
|
||||
**Rate Limit Configuration**:
|
||||
- Anonymous submission: 10 requests per hour per IP address
|
||||
- Dashboard/admin routes: No rate limiting (authenticated users only)
|
||||
- Rate limit headers included in response: `X-RateLimit-Limit`, `X-RateLimit-Remaining`, `X-RateLimit-Reset`
|
||||
|
||||
**Considerations**:
|
||||
- IP-based limiting can be circumvented via VPN/proxy but sufficient deterrent for casual abuse
|
||||
- Behind proxy/load balancer: Configure Flask-Limiter to read `X-Forwarded-For` header
|
||||
- Note: FR-055 requires no fingerprinting for identification - rate limiting is for abuse prevention only, not user tracking
|
||||
|
||||
**Alternatives Considered**:
|
||||
- **CAPTCHA (hCaptcha/reCAPTCHA)**: Adds friction to user experience, contradicts "lower barriers" goal
|
||||
- **Redis-backed rate limiting**: Unnecessary complexity for single-server MVP
|
||||
- **No rate limiting**: Violates FR-061 requirement, leaves system vulnerable to abuse
|
||||
- **Token bucket per session**: Requires session tracking for anonymous users, violates anonymity
|
||||
|
||||
**Dependencies**:
|
||||
- `Flask-Limiter` extension
|
||||
|
||||
---
|
||||
|
||||
## 6. Flask Best Practices for Simple HTML Applications
|
||||
|
||||
### Decision: Server-side rendering with Jinja2 templates, no JavaScript
|
||||
|
||||
**Rationale**:
|
||||
- User explicitly specifies "plain html using flask" and "does not use any css frameworks or javascript libraries"
|
||||
- Server-side rendering eliminates frontend build complexity
|
||||
- Jinja2 included with Flask, no additional dependencies
|
||||
- Forms use standard HTTP POST/GET, progressive enhancement approach
|
||||
|
||||
**Template Approach**:
|
||||
- Minimal inline CSS for basic layout (no framework)
|
||||
- Semantic HTML5 for accessibility
|
||||
- Server-side form validation with error display
|
||||
- Standard browser form controls (no custom widgets)
|
||||
|
||||
**Form Handling**:
|
||||
- POST requests for submissions
|
||||
- Server-side validation with error messages
|
||||
- Flash messages for user feedback
|
||||
- Redirect-after-POST pattern to prevent duplicate submissions
|
||||
|
||||
**No JavaScript Requirement** (Edge Case from spec.md line 105):
|
||||
- "What happens when a user's browser doesn't support JavaScript"
|
||||
- Answer: Application works fully without JavaScript (no JS used)
|
||||
- File uploads work via standard HTML `<input type="file" multiple>`
|
||||
|
||||
**Best Practices Applied**:
|
||||
- Flask app factory pattern for testability
|
||||
- Blueprint organization for routes
|
||||
- Environment-based configuration
|
||||
- CSRF protection via Flask-WTF (even for simple forms)
|
||||
|
||||
**Dependencies**:
|
||||
- Flask (includes Jinja2)
|
||||
- Flask-WTF for CSRF protection
|
||||
|
||||
---
|
||||
|
||||
## 7. Python Dependency Management
|
||||
|
||||
### Decision: requirements.txt with pinned versions for reproducibility
|
||||
|
||||
**Rationale**:
|
||||
- Simplest dependency management for Flask application
|
||||
- No need for Poetry/Pipenv complexity in MVP
|
||||
- Pin exact versions for reproducibility
|
||||
- Virtual environment assumed for isolation
|
||||
|
||||
**Core Dependencies** (estimated):
|
||||
```
|
||||
Flask==3.0.0
|
||||
Flask-Login==0.6.3
|
||||
Flask-Limiter==3.5.0
|
||||
Flask-WTF==1.2.1
|
||||
anthropic==0.8.0
|
||||
clamd==1.0.2
|
||||
bcrypt==4.1.2
|
||||
PyYAML==6.0.1
|
||||
pytest==7.4.3
|
||||
pytest-flask==1.3.0
|
||||
```
|
||||
|
||||
**Development Dependencies**:
|
||||
- pytest, pytest-flask for testing
|
||||
- black for code formatting
|
||||
- flake8 for linting
|
||||
|
||||
---
|
||||
|
||||
## Technology Stack Summary
|
||||
|
||||
| Component | Technology | Rationale |
|
||||
|-----------|-----------|-----------|
|
||||
| **Web Framework** | Flask 3.0+ | Lightweight, simple, widely supported |
|
||||
| **Template Engine** | Jinja2 (built-in) | Server-side rendering, no JS needed |
|
||||
| **AI Provider** | Claude API (Anthropic) | User-specified, abstracted for future flexibility |
|
||||
| **Authentication** | Flask-Login + bcrypt | Standard session management, secure passwords |
|
||||
| **Rate Limiting** | Flask-Limiter | Prevent abuse, simple IP-based approach |
|
||||
| **Malware Scanning** | ClamAV + clamd | Open-source, reliable, on-premise |
|
||||
| **Storage** | Filesystem (YAML + Markdown) | Matches user requirements, simple, no DB |
|
||||
| **Testing** | pytest + pytest-flask | Industry standard, good Flask integration |
|
||||
| **Python Version** | 3.11+ | Modern, stable, good performance |
|
||||
| **Deployment** | Gunicorn (WSGI) | Production-ready Flask server |
|
||||
|
||||
---
|
||||
|
||||
## Non-Functional Requirements Research
|
||||
|
||||
### Performance Considerations
|
||||
|
||||
**Concurrent Submissions** (SC-012):
|
||||
- Target: 100 concurrent submissions without errors
|
||||
- Flask + Gunicorn with 4-8 worker processes should handle this
|
||||
- File I/O is bottleneck: consider async I/O if performance issues arise
|
||||
- AI analysis happens asynchronously (background task) to not block submission response
|
||||
|
||||
**Dashboard Performance** (SC-008):
|
||||
- Target: Load 1000 items in <3 seconds
|
||||
- File-based approach: Index product feedback directories, cache counts
|
||||
- Implement pagination (50 items per page)
|
||||
- Use lazy loading for file attachments (links, not embedded content)
|
||||
|
||||
**AI Analysis Time** (SC-007):
|
||||
- Target: <30 seconds for 95% of submissions
|
||||
- Claude Haiku model typically responds in 2-5 seconds for translation/categorization
|
||||
- Timeout: 45 seconds before marking as failed
|
||||
- Queue-based processing if needed (Python `queue` module or simple file-based queue)
|
||||
|
||||
### Security Considerations
|
||||
|
||||
**Anonymity Enforcement** (FR-055, SC-010):
|
||||
- Do NOT log IP addresses in feedback metadata
|
||||
- Rate limiting uses IP for abuse prevention only, not stored with feedback
|
||||
- No session cookies for anonymous submission
|
||||
- No analytics/tracking scripts
|
||||
|
||||
**File Upload Security**:
|
||||
- Validate MIME types server-side (don't trust client)
|
||||
- ClamAV scanning before storage
|
||||
- Sanitize filenames to prevent directory traversal
|
||||
- Store outside web root, serve via Flask route with access control
|
||||
|
||||
**HTTPS Requirement** (FR-064):
|
||||
- Deployment guide must specify reverse proxy (nginx) with TLS
|
||||
- Redirect HTTP to HTTPS
|
||||
- HSTS headers recommended
|
||||
|
||||
---
|
||||
|
||||
## Open Questions for Implementation Phase
|
||||
|
||||
1. **Asynchronous AI Analysis**: Should analysis happen synchronously (user waits) or asynchronously (background job)?
|
||||
- **Recommendation**: Asynchronous - return success immediately, process in background
|
||||
- Implement simple file-based queue or use Python `threading` for MVP
|
||||
|
||||
2. **Admin Bootstrap**: How is the first admin user created?
|
||||
- **Recommendation**: CLI command or config file initialization script
|
||||
|
||||
3. **Email Notifications**: Out of scope (line 255) but commonly requested
|
||||
- **Recommendation**: Document as future enhancement, design hooks for extensibility
|
||||
|
||||
4. **Logging Strategy**: Structured logs for operational monitoring?
|
||||
- **Recommendation**: Python `logging` module, JSON format, separate file per environment
|
||||
|
||||
---
|
||||
|
||||
## Research Validation
|
||||
|
||||
All NEEDS CLARIFICATION items from Technical Context have been resolved:
|
||||
|
||||
| Item | Resolution | Document Section |
|
||||
|------|------------|------------------|
|
||||
| AI Integration | Pluggable interface, Claude as default | §1 |
|
||||
| File Upload - Malware Scanning | ClamAV + clamd | §2 |
|
||||
| File Upload - Storage Location | Filesystem: `data/products/.../feedback/.../attachments/` | §3 |
|
||||
| Authentication | Flask-Login + server-side sessions + bcrypt | §4 |
|
||||
| Rate Limiting | Flask-Limiter, 10/hour/IP | §5 |
|
||||
|
||||
**Next Phase**: Proceed to Phase 1 (data-model.md, contracts, quickstart.md)
|
||||
Reference in New Issue
Block a user