Add comprehensive post-processing features to convert downloaded Excel files into structured data and merge them into a single output file. New modules: - extractor_core.py: Stateless pure functions for web operations - excel_converter.py: Excel to DataFrame conversion utility - tests/test_extractor_real.py: Real data extraction test suite Enhanced features: - post_process_downloads(): Convert and merge multiple Excel files - extract_and_process(): Complete workflow in single call - cleanup_temp_files(): Optional cleanup of temporary downloaded files - Field name mapping for standardized output columns Dependencies: - pandas>=2.0.0 for data manipulation - openpyxl>=3.1.0 for Excel file handling Documentation: - Updated CLAUDE.md with new module references - Added API documentation for extractor components Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
164 lines
8.0 KiB
Markdown
164 lines
8.0 KiB
Markdown
# CLAUDE.md
|
|
|
|
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
|
|
|
|
## Project Overview
|
|
|
|
BIPAuto is a Python automation framework for interacting with Yonyou BIP (用友BIP) enterprise resource planning systems. The project uses Playwright for browser automation to perform login, logout, and other operations on the web-based ERP system.
|
|
|
|
## Development Environment
|
|
|
|
**Virtual Environment**: All Python scripts must be executed within the local `.venv` virtual environment.
|
|
|
|
```bash
|
|
# Create virtual environment (if not exists)
|
|
python -m venv .venv
|
|
|
|
# Activate (Windows Git Bash)
|
|
source .venv/Scripts/activate
|
|
|
|
# Install/update dependencies
|
|
pip install -r requirements.txt
|
|
```
|
|
|
|
## Environment Configuration
|
|
|
|
The project relies heavily on environment variables loaded from `.env` file in the project root. Never commit `.env` to version control - use `.env.example` as a template.
|
|
|
|
**Critical Environment Variables:**
|
|
- `PLAYWRIGHT_BROWSERS_PATH` - Path to Playwright browser installation (custom path, not default)
|
|
- `ERP_URL` - Base URL for the ERP system
|
|
- `ERP_USERNAME` - Login username
|
|
- `ERP_PASSWORD` - Login password
|
|
- `ERP_HEADLESS` - Whether to run browser in headless mode (true/false)
|
|
- `ERP_IGNORE_HTTPS_ERRORS` - Whether to ignore HTTPS certificate errors (true/false)
|
|
|
|
Note: Test scripts are responsible for loading environment variables and passing configuration to utility functions.
|
|
|
|
## Architecture
|
|
|
|
### Module Structure
|
|
|
|
**`utils/auth.py`** - Core authentication module (pure functions)
|
|
- `login(playwright, username, password, url, headless, ignore_https_errors, verbose=True)` - Handles Yonyou BIP login with automatic force-login popup detection. Requires all parameters (playwright, username, password, url, headless, ignore_https_errors).
|
|
- `logout(page)` - Performs logout with confirmation dialog handling.
|
|
- Returns tuple: `(browser, context, page, main_frame)` where `main_frame` is the forwardFrame iframe.
|
|
- Callers are responsible for browser lifecycle management (context.close(), browser.close()).
|
|
|
|
**`utils/discrete_material_plan/extractor_core.py`** - Core web operations module (pure functions)
|
|
- `navigate_to_discrete_material_page(main_frame, page)` - Navigate to discrete material plan maintenance page. Returns `(work_frame, page1)`.
|
|
- `setup_query_interface(work_frame)` - Initialize query interface by selecting order number query tab and setting page size.
|
|
- `fill_and_search_orders(work_frame, order_ids)` - Fill order IDs into search textbox and trigger search.
|
|
- `download_batch_data(work_frame, page, order_ids, batch_index, download_dir)` - Execute download workflow for a single batch. Returns downloaded file path.
|
|
- `execute_batch_download_workflow(work_frame, page, order_ids, batch_index, download_dir)` - Complete workflow: fill orders, search, and download. Returns downloaded file path.
|
|
- All functions are stateless and accept required parameters explicitly. No logging or progress reporting - callers handle that.
|
|
- **See**: `docs/discrete_material_plan_extractor_core.md` for complete API documentation and usage examples.
|
|
|
|
**`utils/discrete_material_plan/extractor.py`** - High-level extractor component
|
|
- `DiscreteMaterialPlanExtractor` class - Batch processing wrapper with progress reporting and error handling.
|
|
- `extract_from_file()` function - Convenience function to extract data from order IDs in a file.
|
|
- **New Methods**:
|
|
- `post_process_downloads()` - Convert and merge downloaded Excel files into structured format
|
|
- `extract_and_process()` - Complete workflow: extract data and post-process in one call
|
|
- Accepts ID list and file paths as parameters, handles session management, batch processing, Excel conversion, and file merging.
|
|
- **Documentation**:
|
|
- Basic Usage: `docs/discrete_material_plan_extractor.md` - Basic API and usage examples
|
|
- Component Guide: `docs/extractor_component_guide.md` - Complete guide with Mermaid diagrams
|
|
- Post-Processing: `docs/extractor_post_processing.md` - Excel conversion and merging features
|
|
|
|
**`utils/discrete_material_plan/excel_converter.py`** - Excel data conversion utility
|
|
- `ExcelConverter` class - Converts raw Excel reports to structured DataFrames.
|
|
- `convert(input_file, output_file)` - Converts Excel file and returns DataFrame.
|
|
- Handles nested order structures, flattens to table format, applies field name mapping.
|
|
- Used internally by extractor for post-processing, can also be used standalone.
|
|
- **See**: `docs/extractor_post_processing.md` for usage examples.
|
|
|
|
**`tests/`** - Test suite
|
|
- All test files must add `PROJECT_ROOT` to `sys.path` to import `utils` modules
|
|
- Tests follow pattern: load dotenv, set browser path environment variable, run Playwright operations
|
|
- All output and log messages use English only (Chinese text reserved for page element selectors)
|
|
|
|
### Page Interaction Pattern
|
|
|
|
Yonyou BIP uses a nested iframe structure:
|
|
1. Main page contains `#forwardFrame` iframe
|
|
2. All form elements (username, password, buttons) are located within this iframe
|
|
3. Use `page.locator("#forwardFrame").content_frame` to access the iframe
|
|
4. Element selectors use role-based locators with Chinese names matching the UI:
|
|
- `get_by_role("textbox", name="用户名")` for username field
|
|
- `get_by_role("button", name="登录")` for login button
|
|
|
|
### Force Login Handling
|
|
|
|
The system automatically detects and handles a force-login confirmation dialog that appears after clicking the login button. The code checks for a "确定" (Confirm) button and clicks it if present.
|
|
|
|
### URL Construction Pattern
|
|
|
|
Callers must construct the complete login URL before passing to `login()`:
|
|
|
|
```python
|
|
url = f"{os.getenv('ERP_URL').rstrip('/')}/yonbip/resources/uap/rbac/login/main/index.html"
|
|
```
|
|
|
|
### Demo Data
|
|
|
|
**`id-demo.txt`** - Demo production IDs for testing
|
|
- Contains 15 sample production IDs (SC70202603240xxx format)
|
|
- Used by test scripts and demos for web operations testing
|
|
- Format: One ID per line, plain text file
|
|
|
|
## Common Commands
|
|
|
|
```bash
|
|
# Run authentication configuration test
|
|
python tests/test_auth_config.py
|
|
|
|
# Run login/logout test
|
|
python tests/test_login.py
|
|
|
|
# Run web operations tests
|
|
python tests/test_web_operations.py
|
|
|
|
# Run extractor component tests
|
|
python tests/test_extractor_component.py
|
|
|
|
# Run real data extraction test (uses tests/id-demo.txt)
|
|
python tests/test_extractor_real.py
|
|
|
|
# Or use quick run scripts
|
|
./run_extractor_test.bat # Windows
|
|
./run_extractor_test.sh # Linux/Mac
|
|
|
|
# Run web operations demo
|
|
python tests/demo_web_operations.py
|
|
|
|
# Run extractor component demo
|
|
python demo_extractor_component.py
|
|
|
|
# Run any test with virtual environment
|
|
source .venv/Scripts/activate && python tests/<test_file>.py
|
|
```
|
|
|
|
## Documentation
|
|
|
|
**`docs/INDEX.md`** - Complete documentation index
|
|
- Quick start guide, API reference, examples, and troubleshooting
|
|
- Central hub for all documentation
|
|
|
|
**`docs/extractor_component_guide.md`** - Complete component guide with Mermaid diagrams
|
|
- Architecture overview, class structure, workflow diagrams
|
|
- State management, error handling, integration examples
|
|
- Visual documentation using Mermaid graphs
|
|
|
|
## Code Conventions
|
|
|
|
1. **English Only**: All user-facing output, docstrings, comments, and log messages must be in English. Chinese text is only used for Playwright element selectors matching the actual UI.
|
|
|
|
2. **Environment-First**: Test scripts load environment variables and explicitly pass configuration to utility functions. No hardcoded values in code.
|
|
|
|
3. **Error Handling**: Functions that depend on environment variables should raise clear `ValueError` exceptions when required variables are missing.
|
|
|
|
4. **Path Handling**: Use `pathlib.Path` for all file system operations. Project root is determined as `Path(__file__).resolve().parent.parent` from within `utils/` modules.
|
|
|
|
5. **Test Coverage**: **ALL new components MUST include corresponding unit tests.** When creating new functionality in `utils/` or any other module, you must simultaneously create a test file in `tests/` directory that validates the component's behavior. Tests should verify both success and failure scenarios.
|