# Project Directive: Lead Scraper (Project Name: ArbitrageIQ)

## 1. Context & Purpose
ArbitrageIQ is a CLI-driven automated real-estate lead scraper and AI recommendation engine built for **Nomadic Crucial** and **Reservation Resources**. 
The system extracts real estate listings from third-party platforms (Facebook Marketplace, SpareRoom, Furnished Finder), standardizes the data, categorizes the property, and recommends one of three monetization pathways:
1. **Nomadic Crucial Feed:** Monetize raw lead data for budget NYC renters.
2. **Direct Operator / Master Lease:** Rent/Purchase directly with an option to buy for internal management.
3. **Partner Network:** Invite landlords to list directly on the Reservation Resources platform (Airbnb/Booking.com style model).

---

## 2. Technical Stack
- **Language:** TypeScript (Executed via `tsx` runner for zero-build efficiency)
- **Runtime:** Node.js (v20+)
- **Database & Storage:** Google Firebase (Firestore for data, Firebase Storage for downloaded images)
- **CLI Framework:** Commander.js
- **CLI Rendering Engine:** `@b9g/termdom` (HTML-like progress bars, spinners, and terminal UI)
- **Proxy & Evasion:** BrightData Residential/Data-center Proxies with fingerprint management
- **Logging:** Winston or Pino

---

## 3. Strict Development Rules & Constraints
1. **Git Control:** DO NOT stage (`git add`) or commit (`git commit`) any files unless explicitly requested by the developer in the prompt.
2. **Code Style:** Functional, modular Node.js/TypeScript ES modules (`import/export`). Clear separation of concerns between scrapers, UI, logging, and database operations.
3. **Type Safety:** Maintain strict TypeScript interfaces for all raw and scraped data entities to prevent runtime errors on null/optional listing fields.
4. **Environment Security:** Never hardcode credentials, Firebase keys, or BrightData proxies. Use `.env` files exclusively.
5. **Proxy Optimization (Direct Media Downloads):** Scrapers must only use BrightData proxies for initial DOM navigation, page interaction, and metadata/image URL extraction. All binary image downloads (and uploading to Firebase Storage) MUST be executed outside of the proxy pipeline using direct standard HTTP requests (`fetch`/`axios`) to avoid unnecessary proxy bandwidth costs.

---

## 4. Logging & Observability Requirements
A centralized logging service (`/config/logger.ts`) must support environments (`development` vs `production`) with stdout and file appenders:

- **DEBUG:** Scraper DOM navigation steps, proxy rotation events, raw network payloads.
- **INFO:** Successful listing extractions, CLI step completions, Firebase write confirmations, generated outreach links.
- **WARN:** Anti-bot captcha detections, missing optional listing fields (e.g., landlord name missing), proxy retry attempts.
- **ERROR:** Uncaught exceptions, failed Firebase writes, DOM parsing failures (selector changes), BrightData authentication errors.

All log output must be JSON-formatted in production and colorized readable text in local development environments.

---

## 5. Schema & Data Requirements
Every scraped entity saved to Firestore must follow this  schema structure. This is just a starting point. We may extend or modify if necessary:

```typescript
export enum RecommendationType {
  NOMADIC_CRUCIAL = 'NOMADIC_CRUCIAL',
  MASTER_LEASE = 'MASTER_LEASE',
  PARTNER_NETWORK = 'PARTNER_NETWORK'
}

export interface LandlordInfo {
  name?: string;
  contact?: string;
  profileUrl?: string;
}

export interface Listing {
  id?: string;
  source: 'facebook_marketplace' | 'spareroom' | 'furnished_finder';
  originalUrl: string;
  scrapedAt: string;
  title: string;
  price: number;
  currency: string;
  location: string;
  description: string;
  images: string[];
  landlord: LandlordInfo;
  recommendation: {
    action: RecommendationType;
    reasoning: string;
    categorizedAt: string;
  };
  outreach: {
    suggestedTemplate: string;
    trackedLink: string;
  };
  analytics: {
    linkClicks: number;
    converted: boolean;
    status: 'NEW' | 'CONTACTED' | 'CONVERTED' | 'ARCHIVED';
  };
}
```



## 6. Directory Structure & Architecture

Image Scraping Pipeline Note
 - **Metadata Extraction:** Performed via scraper (/src/scrapers/) through BrightData proxy.

 - **Image Downloader:** Handled separately via media service (/src/services/imageDownloader.ts) using standard direct HTTP connections without proxy agent configurations.

### Project Directory Layout (ArbitrageIQ)


```bash
/arbitrage-iq
├── bin/
│   └── cli.ts              # Commander.js entry point & termdom UI setup
├── config/
│   ├── firebase.ts         # Firebase DB & Storage initialization
│   └── logger.ts           # Logging setup
├── src/
│   ├── types/
│   │   └── listing.ts      # TypeScript Interfaces & Enums
│   ├── scrapers/
│   │   ├── baseScraper.ts  # Parent scraper with BrightData proxy logic
│   │   └── facebook.ts     # FB Marketplace scraper implementation
│   ├── services/
│   │   ├── imageDownloader.ts # Bypasses proxy to fetch & upload raw image streams
│   │   ├── categorizer.ts  # Recommendation/categorization engine
│   │   ├── analytics.ts    # Link tracking generator
│   │   └── outreach.ts     # Outreach copy generator
├── GEMINI.md               # Context & directives for LLM collaborators
├── tsconfig.json           # Minimal TypeScript configuration
├── package.json
└── .env.example
```




## 7. External Resources & Official Documentation
- **Bright Data Facebook Marketplace API:** https://docs.brightdata.com/api-reference/scrapers/social-media-apis/facebook-marketplace-collect-by-url
- **Firebase JS Reference API:** https://firebase.google.com/docs/reference/js



## 8. Versioning & Package Management Rules
1. **Semantic Versioning (SemVer):** The project follows strict Semantic Versioning (`MAJOR.MINOR.PATCH`):
   - **MAJOR:** Breaking changes in scraper outputs, database schemas, or CLI interface options.
   - **MINOR:** Backward-compatible new features (e.g., adding a new scraper source or categorization rule).
   - **PATCH:** Backward-compatible bug fixes, DOM parsing updates, or logging adjustments.
2. **Package Synchronization:** Whenever version increments occur, the version declared in `package.json` must strictly match the CLI metadata output (`cli.version(...)`) and application runtime logs.
3. **Dependency Integrity:** Lockfiles (`package-lock.json`) must be checked into version control. Upgrading core dependencies (such as Commander.js, Firebase, or termdom) must be explicitly tested before bumping version numbers.