# HBL Multi-Stage Product Scraper

Local Laravel application for discovering products from the public HBL installment catalog, selecting products for detailed scraping, and downloading product images with concurrent queue workers.

## Requirements

- PHP 8.2+ with `pdo_mysql`, `curl`, `mbstring`, and `openssl`
- MySQL 8+
- Node.js 20+ and Google Chrome/Chromium
- Composer (a local `composer.phar` can be used)

## Installation

```bash
php composer.phar install
npm install
cp .env.example .env
php artisan key:generate
mysql -uroot -e "CREATE DATABASE IF NOT EXISTS snpl_scraper CHARACTER SET utf8mb4 COLLATE utf8mb4_unicode_ci"
php artisan migrate
php artisan storage:link
```

Set the correct MySQL username/password in `.env`. If system Chrome is unavailable, install Playwright's browser:

```bash
npx playwright install chromium
```

## Running

Run the web application:

```bash
php artisan serve
```

In a second terminal, start the configured worker pool (four workers by default):

```bash
php artisan scraper:workers
```

Open `http://127.0.0.1:8000`, run **Discover Products**, then use the product inventory checkboxes to scrape details or download images.

## Configuration

The `.env` scraper settings control the source URL, worker count, delay, timeout, retries, Node binary, and image disk. `SCRAPER_WORKERS` controls concurrency; `SCRAPER_DELAY_MS` applies a per-item throttle. Every product job has independent attempts and error state and can be retried from the dashboard.

Images are stored under `storage/app/public/products/{product_id}` and exposed through Laravel's public storage link. The raw source payload is retained alongside normalized fields so extraction mappings can be updated without losing source data.

## Tests

```bash
php artisan test
```

Tests use in-memory SQLite and do not contact the live source.
