Beautifulsoup jobs
Website Content Scraping Required I need content to be extracted from a website and organized in a structured format. The task includes scraping text, images (if required), and other relevant information while maintaining accuracy and proper formatting. Requirements...maintaining accuracy and proper formatting. Requirements: - Extract content from the specified website. - Preserve headings, paragraphs, and content structure. - Organize the extracted data in Excel, CSV, or Word (as required). - Ensure the data is clean, complete, and free from duplicates. - Deliver the project within the agreed timeline. Experience with web scraping tools (such as Python, BeautifulSoup, Scrapy, Selenium, or similar) is preferred. Please mention your approach, estimated timeline, and cost in your...
...Announcement, Specialty, Experience, Experience, Book on, summary, badges and designations, (Clinic schedule, Fee, Clinic Name, Clinic Location,) of all clinics they have. Education, Med School, Residency, Fellowship Training, Certifications, online clinic hours and availability, Online clinic fee, Affiliations You may harvest the information with the tooling of your choice—Python (BeautifulSoup, Scrapy, Selenium), R, or another reliable stack—so long as the final file imports seamlessly. I am flexible on the exact format (Excel, CSV, or database dump); let me know what works best for your workflow and I’ll confirm before we start. Accuracy is critical. I will verify that every profile on the site is represented and that each required field is populat...
Project Overview I have a digital corporate directory database containing 4,100+ company profiles. The database layout is locked. I need a technical data specialist to extract this information via data-scraping scripts or OCR tools into a cleanly formatted, sortable Microsoft Excel sheet. Project Requirements & Rules No Manual Typing: You must use automated scraping scripts (Python/Selenium/BeautifulSoup) or premium OCR parsers to prevent human entry errors in phone numbers and email addresses. Zero Data Deletion: Do not delete any rows, states, or categories. Keep the entire dataset 100% intact. Delivery Timeframe: 48 hours from project award. Final Excel Grid Deliverables The final .xlsx sheet must feature standard multi-column filters on the top row, mapped to these exact...
Title: Build a Local B2C Homeowner Lead Alert Tool (Fixed Budget: $200) Project Overview: I am a roof restoration and rejuvena...the post. 7. Output Deliverable: A simple local Python script that saves these results cleanly into a local Excel/CSV file on my computer. A plain text file () with simple, step-by-step installation instructions for a non-technical user must be included. No external databases or cloud hosting required. Required Skills: - Python - Web Scraping / Automation (Selenium, Playwright, or BeautifulSoup) - Social Media / Public Forum Data Extraction To apply, please start your message with the word "REJUVENATE" so I know you read the description. Briefly explain what framework you would use to build this and confirm it can run locally without recurring...
...experienced web-scraping specialist who can reliably pull data from Facebook and Instagram. The focus is on public, real-time content; I want clean, structured results that can be analysed right away without additional tidying on my side. To be sure your approach works, please show me a short sample first—10-20 recent records from each platform will do. I am happy with any proven stack (Python + BeautifulSoup/Scrapy, Node + Puppeteer, Selenium, API work-arounds, etc.) as long as you can demonstrate stability, speed, and respect for each platform’s rate limits. Deliverables (all items required): • One working script or tool that scrapes Facebook and Instagram as agreed • A sample dataset for my review before full engagement • Clear instructions or ...
1. Project Title () Development of an Automated NGO Darpan Data Extraction, Processing, and Management System Using Python, Selenium, BeautifulSoup, SQLite, and Streamlit ________________________________________ 2. Abstract The NGO Darpan portal is India's central repository for registered Non-Governmental Organizations (NGOs), Voluntary Organizations (VOs), Trusts, Societies, and Section 8 Companies. The portal contains publicly available organizational information including registration details, addresses, office bearers, operational sectors, and contact information. The objective of this project is to develop a Python-based automation system capable of systematically collecting publicly available NGO information after authorized user interaction where
...authentication, API keys, rate limits, retries, and logging. * Build error handling for failed/partial imports. * Maintain source attribution for every imported field. * Ensure no duplicate horse/race records are created. * Add tests for every new integration. ## Required Skills * Python * FastAPI * PostgreSQL * SQLAlchemy * REST API integration * JSON/XML/CSV data mapping * HTML parsing with BeautifulSoup/lxml * Data validation * API authentication and rate-limit handling * Git ## Preferred Skills * Sports data APIs * Horse racing or betting data experience * Web scraping/parser experience * PDF/OCR data extraction * Background jobs / schedulers * Docker ## Important Requirements * Do not redesign the existing backend. * Do not create duplicate database tables unless requi...
...website and would like the final dataset delivered as a CSV. The CSV must include every visible product together with the key details you typically find on a catalogue page—name, price, image URL, description, SKU or ID, and category path. I do not need user reviews, contact details, or any ongoing scheduling; this is strictly a single execution job. Your script can be in Python (requests + BeautifulSoup, Scrapy, or Selenium if the site is JavaScript-heavy) or another language you are comfortable with, as long as it runs reliably on a standard desktop environment. Deliverables: • The CSV file with all product rows and clearly labeled columns. • The executable script or notebook, well-commented so I can rerun it later if the site structure remains unchang...
...looking for a detail-oriented web-scraping specialist who can pull data from virtually any public-facing website and hand it back to me neatly organised in an Excel workbook. The site type and data fields will vary from project to project—sometimes it might be product listings, other times articles, reviews, or contact details—so adaptability and solid experience with tools such as Python (BeautifulSoup, Scrapy, Selenium) or equivalent are essential. What matters most is accuracy, clean formatting, and a repeatable process I can rerun in the future. Along with the finished .xlsx file, please include either the script or clear documentation of your method so I can update the scrape if the source site changes. If you have questions about pagination, login barriers, or...
...CRM • After export, the same script (or a companion script) uploads both CSVs to the relevant Supabase tables. • On insert, each record must automatically receive an assigned salesperson ID (simple round-robin logic is fine). • A price field should be calculated from square metres already present in the listing and stored alongside the record. Preferred stack: Python 3 with BeautifulSoup or Playwright for scraping and the official Supabase client for the upload, but I’m open to Node.js if you have a stronger approach. Deliverables • Clean, well-commented source code. • Two sample CSVs generated from a test URL. • Brief README covering setup, environment variables, and how to run both the scrape and the upload. Once...
I'm seeking a skilled Python developer to work on a project. The specific type of project, primary function, and preferred libraries/frameworks are currently undecided. Ideal Skills and Experience: - Proficiency in Python - Experience with web applications, data analysis, or automation - Familiarity with Django/Flask, Pandas/Numpy, or Scrapy/BeautifulSoup - Strong problem-solving skills - Ability to work independently and meet deadlines Please provide relevant experience and a brief project approach in your bids.
...Python-based project that will either evolve into a lightweight FastAPI service or a robust web-scraping pipeline—whichever proves the better fit once we start prototyping. Clean, asynchronous code, thoughtful error handling and respect for rate limits are non-negotiable, whether you lean on FastAPI’s dependency-injection patterns or a scraping stack such as Scrapy, Playwright, Selenium, or BeautifulSoup. Before we dive into technical details, I’d like to see concrete proof of your expertise. Please share past work: a GitHub repo, a running demo, or any code snippets that highlight your mastery of FastAPI endpoints, background tasks, pydantic models, or large-scale data extraction from dynamic sites. Real examples will help me understand your coding style and d...
...offerings across the sites. Please deliver the data in a single Google Sheets workbook, with a dedicated tab for each source site and clear column headers. So I can plan my budget, let me know your price per site along with an estimated turnaround time once you see the list and any anti-scraping measures that might require work-arounds. • If you already have tooling or scripts in Python, BeautifulSoup, Scrapy, Selenium, or similar, feel free to mention it—speed and reliability matter more to me than the specific stack....
...content outline * Priority (High / Medium / Low) * Estimated SEO impact ### 5. Reporting Generate a structured report (Markdown or JSON) containing: * Executive summary * Competitor strengths * Our content weaknesses * Prioritized recommendations * Suggested content outline * Overall SEO score for both pages ## Technical Requirements Preferred technologies: * Python * Playwright or Selenium * BeautifulSoup * OpenAI, Claude, or Gemini APIs * LangGraph, LangChain, CrewAI, or similar AI agent frameworks (optional) ## Nice to Have * Strong understanding of SEO and content optimization * Experience building AI agents or automated research tools * Experience with semantic SEO and topical authority * Previous work in real estate or marketplace websites ## When Applying Please include: *...
...Extract headline, key statistics, and the article link. • Pre-append a short personalized note that I can define in a config file or an Airtable/Google Sheet cell before each post goes live. • Push the final content to my LinkedIn company page through the LinkedIn API, respecting posting limits and avoiding duplication. I’m comfortable running the solution on a small VPS, so Python with BeautifulSoup/Playwright, Zapier, , or any practical stack is fine—just keep the setup straightforward and well-documented. Deliverables 1. Source code or no-code scenario file. 2. Read-me / screen-share walkthrough showing how to swap in new UAE news sites and edit the personalized message. 3. A quick test proving one live post on my LinkedIn page. Once everythi...
...matching data directly from a company’s website when that adds value. Core needs • A script, API, or lightweight app that inputs a keyword or industry and returns clean profile data (name, role, company, public URL). • Smart filtering to remove duplicates and obvious non-prospects. • Export options—CSV or JSON at minimum—for easy hand-off to marketing systems. Tech is up to you; Python (BeautifulSoup, Selenium, Scrapy) or a comparable stack is welcome as long as it respects LinkedIn’s limits and complies with website terms. Please send a gedetailleerd projectvoorstel outlining: 1. Your approach to bypassing anti-scraping measures without violating TOS. 2. Key milestones from prototype to final delivery. 3. Examples of similar scra...
I need a Python automation script that runs sm...simple system operations, or another Python-friendly task we identify together. Once we agree on the exact routine, I will share all supporting files, URLs, or process notes. Deliverables • A clean, well-commented .py script (Python 3.x) ready to execute on Windows without manual tweaks • A concise README with setup instructions and a list of required libraries (e.g., pandas, requests, BeautifulSoup, pyautogui, etc.) • Meaningful logging and graceful error handling so I can track what the script is doing and why I value straightforward code I can maintain myself later, so please keep the structure logical and modular. Let me know your proposed approach, turnaround time, and any clarifying questions so we can g...
I have a stream of numerical information being pulled automatically from several news sites, and I need that raw output cleaned,...routine that parses each scrape, identifies the key figures buried in the articles, applies consistent codes to them, and drops everything into a structured file (CSV or JSON works for me). You’ll receive: • the current scraping script and a sample of the raw dump • a field dictionary showing how each number should be labeled or categorised I’m expecting your returned script (Python preferred—BeautifulSoup/Scrapy plus pandas is perfect, but use what you like) along with the final processed dataset and a brief read-me so I can rerun or extend the pipeline later. Accuracy of the coding and reproducibility of results will be ...
...and sub-sections, but the structure is consistent and does not change. I would like everything that appears in those tabs captured—no filtering, no summarising, simply all available fields. Key preferences • Output: one clean Excel workbook (.xlsx) per run. • Frequency: I plan to run the job once a month, so the solution should be fire-and-forget—ideally a Python script using requests/BeautifulSoup, Selenium, or an equivalent approach that I can trigger manually. • Budget: I have about $20 in mind but I’m open to reasonable proposals if you can demonstrate reliability and speed. Deliverables 1. Fully commented source code or macro that fetches the data and produces the Excel file. 2. A sample run showing today’s extraction. 3. ...
...should be reliable, detail-oriented, and capable of handling various Python-related tasks independently. Requirements: Strong knowledge of Python Experience with Web Scraping and Data Extraction API Integration and Automation Data Processing and Data Management Ability to troubleshoot and optimize existing scripts Good communication skills Ability to meet deadlines Preferred Skills: Selenium BeautifulSoup Scrapy Pandas Requests Flask or Django (optional) Database experience (MySQL, PostgreSQL, SQLite) What I Offer: Long-term work opportunities Multiple projects every month Clear requirements and communication Prompt payment for completed work Please include: Your Python experience Examples of previous projects Your hourly rate or fixed-price expectations Your availability...
...AppleWebKit/537.36' } bulk_data = [] for url in URLS_TO_CRAWL: try: print(f"Crawling: {url}") response = (url, headers=headers, timeout=10) if response.status_code != 200: print(f"Skipping {url}: Status code {response.status_code}") continue soup = BeautifulSoup(, '') # குறிப்பிட்ட HTML கூறுகளை மட்டும் எடுத்தல் (Headings, Paragraphs, Custom tags) content_pieces = [] # h1, h2, p மற்றும் உங்களின் பிரத்யேக Tag-களை இங்கே குறிப்பிடவும் for element in soup.find_all(['h1', 'h2', 'p', 'custom-tag']): ...
...straightforward: crawl each page, parse the HTML for the specific elements I’ll identify (headings, paragraphs, and a couple of custom tags), normalise any odd characters, then bulk-insert the results so the database is immediately query-able. A repeatable solution matters because I’ll be running the same process weekly as the sites update. I’m comfortable if you build the scraper in Python—BeautifulSoup, Scrapy, or Selenium are all fine—or you can propose another language or library you prefer, as long as it reliably handles pagination and throttles requests to stay respectful of the hosts. Deliverables: • A well-commented script or small codebase that performs the crawl, parse, and SQL insert in one run. • A SQL file (or direct push...
...week, past month, or leave it open) • Work type (full-time, part-time, remote—whatever option I pass in) For every listing that matches those filters, the scraper must return a structured record containing: • Job link • Job title • Company name • Full job description • Salary information (when LinkedIn shows it) • Experience level • Job type Stack is your choice—Python with BeautifulSoup/Selenium, Node with Puppeteer, or any other robust approach—as long as the final solution: 1. Runs from the command line with a single command. 2. Accepts the three parameters above without code edits. 3. Outputs clean CSV and JSON files in the working directory. Deliverables 1. Fully commented source code. 2. ...
I need every article on my Arabic-language website copied out of its static HTML pages and delivered as clean, UTF-8 plain-text files. The site has no dynamic elements or log-in walls, so a straightforward scraper (Python + BeautifulSoup, wget, or any similar tool you prefer) should be enough. Please remove all HTML, inline styles, menus, and ads; keep only the article title, body, sub-headings, and any author/date line that appears inside the article itself. Each piece should be saved as its own .txt file, named after the URL slug or the article title (whichever is easier to automate). Diacritics and right-to-left order must stay intact. Deliverables: • A zipped folder of the plain-text files, one per article • The script or command line you used so I can rerun ...
My day rarely looks the same twice, so I need a flexible side-kick who can switch effortlessly between deep-dive research, fast data scraping, AI-powered task execution, and polished document preparation. One hour you might be pulling product data with Python, BeautifulSoup, or Scrapy; the next you could be steering ChatGPT or another LLM to summarise findings, draft a proposal, or even spin up instructions for a third-party service I delegate to. Core responsibilities • Research – everything from quick market snapshots to more academic or product-level analysis, depending on what the week demands. • Data scraping & cleanup – locating reliable sources, extracting the essentials, and presenting them in clear, reusable formats (CSV, Google Sheets, Airtab...
I need a reliable way to collect genuine U-S-A mobile contacts—each record must include the person’s name, cell number, and ph...site’s TOS. Deliverables • A working web application with secure login and dashboard • Scraper modules for the three platform categories above, with easy ability to add more later • Structured output: name, phone, street address, city, state, ZIP • Documentation covering setup, usage, and adding new sources I’m open to the tech stack you feel most comfortable with—Python (Scrapy/BeautifulSoup), Node.js, or similar—as long as it can handle rotating proxies, CAPTCHA solving, and basic anti-ban tactics. Please outline your proposed approach, past experience with large-scale scraping, and an...
...(names, addresses, SSNs, medical record numbers, account details, etc.). • Focus domains: Healthcare, Finance and Education. • Geographic emphasis: North-American publications at this stage (government portals, open-data sites, public court documents, regulatory disclosures, etc.); other regions may follow once this tranche is complete. • Methods: a mix of automated scraping (Python, BeautifulSoup/Scrapy/Selenium or similar) and classic desk research to reach sources that resist automation. • Output: an organised folder structure plus a spreadsheet/JSON catalog listing document title, source URL, date accessed, domain tag, and a short note of the specific PII fields present. Acceptance criteria 1. Minimum 250 unique documents, balanced across the th...
I need a complete, accurate scrape of every dealership listed at...Full physical address • Contact person’s name • Contact person’s email The finished dataset should be delivered as a clean, de-duplicated CSV file ready for import. Quality is essential: no missing rows, no placeholder values, and emails must be verified as coming from the official source. I will spot-check several records before issuing final approval. If you intend to use Python (BeautifulSoup, Selenium, Scrapy, etc.) or another toolset, that’s fine—just be sure the scraper respects pagination, German characters, and rate limits so nothing is skipped or blocked. The price is fixed; please confirm you can supply the complete CSV with all 19,843 dealerships and the extra...
I need contact details scraped from 1-2 matrimony sites for my matrimony venture. Requirements: - Experience in web scraping - Familiarity with matrimony sites - Ability to deliver data in organized format Ideal Skills: - Proficient in Python or similar languages - Knowledge of scraping tools like BeautifulSoup or Scrapy - Attention to detail and data accuracy Please provide samples of previous work and estimated delivery time.
I need contact details scraped from 1-2 matrimony sites for my matrimony venture. Requirements: - Experience in web scraping - Familiarity with matrimony sites - Ability to deliver data in organized format Ideal Skills: - Proficient in Python or similar languages - Knowledge of scraping tools like BeautifulSoup or Scrapy - Attention to detail and data accuracy Please provide samples of previous work and estimated delivery time.
I nee...SKU, description, availability and any visible category tags—into the right columns. Everything is public and loads in plain HTML, so no login is required, but I do want the pull to be 100 % complete, free of duplicates and ready for quick analysis the moment I open the file. Deliverables • One .xlsx file containing the full product list • The reusable scraping script (Python with BeautifulSoup, Scrapy, or a similar tool) or a clear step-by-step method so I can repeat the extraction later Let me know your estimated turnaround and any potential hurdles you see with pagination or rate limits so we can address them up front. In total it there are 160.000 product entries with 5 specifications that I need in an excel or csv file. So its a quick job ...
### **Project Title:** South Africa Census Data Extraction & Formatting (2011 Household Income by Ward Level) ### **Project Description:** I am seeking a ...600 * R 2 457 601 or more #### **Data Validation Requirement:** * The total row count must match the total number of valid South African wards (approx. 4,460 to 4,468 entries). * Missing data, null values, or unmapped ward boundaries must be clearly flagged as NaN rather than left blank or filled with zeroes. #### **Skills Required:** * Data Extraction / ETL * Web Scraping (Python / BeautifulSoup / Requests) * Excel / CSV Formatting * Experience with South African geographic data (Stats SA / MDB shapefiles) is a major advantage. Please state your estimated turnaround time and your approach to handling the extraction...
Assist with WooCommerce maintenance by troubleshooting stock discrepancies, improving product catalog management, and verifying Stripe payment integration. Create a Python script with BeautifulSoup to extract product-related data from selected websites and generate inventory update recommendations.
...HTML data layer (such as the `__NEXT_DATA__` JSON framework). 4. The bot returns a structured summary back to the user displaying: Property Unit Number, Property Building Name, Land Number, and Room Count. Key Technical Requirements: - Complete proficiency with the Telegram Bot API and Python frameworks (python-telegram-bot or telebot). - Strong web scraping and data extraction experience (BeautifulSoup, JSON parsing, requests). - Knowledge of integrating residential proxies or anti-scraping web unblockers. - Integration of a simple, lightweight SQLite database to cache successful searches. This prevents the bot from making repeat API calls for duplicate links, saving running costs. Budget & Milestones: - I expect this project to take 3 to 5 days. - Please provide an all-i...
...script that pulls specific text content from a particular website each day and appends the results to a clean, well-structured CSV. The crawl has to run automatically on a 24-hour schedule, capture only the text I specify, and overwrite or update the file so I can download it at any time. Please handle the full workflow—from parsing the HTML through a stable method (Python with Requests + BeautifulSoup, Scrapy, or a comparable stack) to setting up the daily trigger (cron job, cloud function, or Windows Task Scheduler). The script must cope gracefully with minor layout changes and alert me if anything breaks. Deliverables • Fully commented source code • One sample CSV generated from a test run • Setup notes so I can deploy the job on my own server A...
I want to move away from copying information by hand and instead capture it automatically from a set of websites I will supply. Your job is to build a reliable automated data-en...macro, or low-maintenance tool that visits each URL, extracts the designated data, and appends it to the spreadsheet without duplicates. • Error handling for missing pages or layout changes so the run doesn’t stop mid-way. • Simple instructions so I can trigger the process myself and adjust the target file path when needed. If you already have experience with web scraping libraries like BeautifulSoup, Selenium, or similar, let me know—speed and accuracy matter to me. I will share the site list and sample output template once we start. Looking forward to a clean, disciplined sol...
...– at minimum I expect company name, full postal address, phone, email, and any registration numbers available on the SOS records. • Data hygiene – deduplicate across sources, normalise addresses, and flag obviously invalid phones or emails. • Delivery – a single, well-structured CSV file ready for immediate import. Alongside the file, include the scripts or notebooks you used (Python + BeautifulSoup/Scrapy/Selenium or similar) and a short README so I can rerun the pipeline later. Acceptance criteria 1. At least 95 % of rows must contain all mandatory fields. 2. No more than 2 % duplicate businesses when matching on name + address. 3. Scripts re-create the same CSV on a fresh machine with only the listed dependencies. If this matches your ...
...name/model Vehicle category/type Daily/weekly/monthly pricing Rental rates and fees Taxes or additional charges (if visible) Availability information Pickup and return locations Promotions or discounts Vehicle specifications/features Images/URLs (if applicable) Any other relevant offer details displayed on the website Technical Requirements Preferred technologies: Python Selenium and/or Playwright BeautifulSoup, Scrapy, or similar libraries are acceptable if needed The scraper should: Navigate dynamically loaded pages Handle pagination and multiple locations Work reliably across all available branches/locations nationwide Be modular, clean, and maintainable Include error handling and logging Avoid duplicate records Generate structured output files automatically Deliverables The f...
I need 3,000 fresh business-directory records pulled into a clean spreadsheet. Each row must include: • Company name • Full address (street, city, state, ZIP) • Primary phone number • Contact email The source is a public online directory; I’ll provide the exact URL and any search parameters as soon as we start. Please use an automated method (Python, BeautifulSoup, Scrapy, Selenium … whatever you prefer) that respects rate limits and captures the data accurately—no missing fields or obvious duplicates. Deliver the final file in CSV or Excel and include the script you used so I can rerun it later if needed. I’ll review random samples against the live site; records must match 98 %+ for acceptance. Let me know your estimated tu...
... search‐results URL and returns a clean dataset ready for analysis. For every listing on the page the script must capture: address, postal code, city, property type, and the listing’s direct link. Once scraped, the data should be saved automatically to both XLSX and CSV formats. Core expectations • Command-line or simple GUI input for the URL • Reliable parsing (requests/BeautifulSoup, Selenium or similar) that copes with pagination and common site layout changes • Output written with pandas or openpyxl so the Excel file is properly formatted • Clear, well-commented code plus a brief README showing how to run the script, set up any virtual environment, and change dependencies if updates its HTML Acceptance criteria 1. Running the script on ...
...tasks as they arise. Typical assignments include writing concise modules, cleaning up legacy scripts, and making sure new code slots in neatly with what’s already there. I’ll supply detailed tickets and unit tests; you return clean, well-commented code that passes those tests. Fluency with core Python 3, Git, and virtual environments is essential. Familiarity with libraries such as Pandas, BeautifulSoup, or Selenium will make your life easier, but an eagerness to learn quickly matters more than any single tool. Turnaround is usually one to three days per ticket. For each task, please deliver: • The updated .py files • A brief README or docstring explaining your approach • Confirmation that all supplied tests pass When you reply, include a short sente...
...tweaks rather than major rewrites) and capable of completing the full crawl again without manual intervention. Data quality is important to me, so please build in duplicate detection, basic normalisation (e.g., consistent postcode and phone formatting) and validation checks that flag obviously invalid or missing fields before exporting. I am flexible on the language and libraries—Python with BeautifulSoup / Selenium or Node with Puppeteer are both fine—so feel free to suggest what will deliver the most reliable results while respecting the terms of each site. Deliverables: • Well-documented script with clear setup instructions and inline comments • A sample CSV/Excel file demonstrating at least a few hundred cleaned records • Brief read-me expla...
I need a clean, well-commented Python script that starts from a single root URL, follows every internal link it finds (no keyword or structural filtering at all), and extracts only the visible text content on each page. The script should rely on mainstream libraries—requests plus BeautifulSoup is fine, but feel free to propose Scrapy or an async stack if it fits better. Please keep the code modular so I can later drop individual functions into a bigger application. Core expectations • Crawl every reachable link within the domain, respecting and an adjustable polite delay. • Skip images, PDFs, or other binary assets; focus strictly on textual information. • Save each page’s URL alongside the extracted text in a single output file (CSV or JSON&mda...
...every Florida-based non-profit that has a publicly available Form 990 on IRS.gov. The task breaks down into two parts: first, a scraper that locates each organisation and downloads the latest Form 990 PDF; second, logic that pulls key details from those filings—most importantly the insurance figure whenever it is $10,000 or higher. My preference is a straightforward Python solution (requests/BeautifulSoup, Selenium only if absolutely necessary) that respects IRS rate-limits, can resume after interruptions, and runs headless on a typical Linux box. Deliverables • Well-commented source code with an easy way to set the filing year and output folder • A folder of PDFs for every qualifying Florida charity found • A CSV summarising: EIN, organisation name, ...
...reputable online directories, and recognised industry publications. Please harvest the data accurately, avoid duplicates, and verify that every record is current. Deliverables • One Excel workbook with separate, clearly labeled sheets (or columns) for company details, product lines, and contact information. • A short note explaining any assumptions, filters, or automated tools/scripts (Python, BeautifulSoup, Selenium, Scrapy, etc.) you used so I can reproduce or update the crawl later. • A quick quality check summary highlighting any gaps you could not fill and why. I’d like the first sample of 25 companies within a few days so we can confirm structure before you proceed to the full scrape. Let me know your estimated turnaround time and any questions yo...
...should be clean, commented, and handed over in full, as I want to maintain or extend it internally. If you have an existing framework or library in Python, Node.js, or another language that can speed development, I’m open to it as long as installation remains straightforward and everything runs locally. Please outline your proposed approach, the tools or APIs you plan to use (e.g., Tweepy, BeautifulSoup, Selenium, etc.), and a rough delivery timeline when you bid....
I need a reliable way to collect product details from several e-commerce sites. I am open to either a fully automated scraper (Python, Selenium, BeautifulSoup, Playwright, etc.) or a well-structured manual process if that proves more stable for the target storefronts. Scope • Target pages: standard product listings and their individual detail pages on the chosen e-commerce sites. • Data fields: everything typically shown on a product page—title, SKU, description, images (URLs are fine), specifications, and the price shown at the moment of capture. • Output: please compile the final dataset in HTML or PDF, whichever reliably preserves the product information and any inline images. If your workflow first generates CSV/JSON/Excel and then converts, that is f...
... Key points you should know: • The site is public but rate-limited, so the script must handle pagination, delays or rotating headers/IPs if necessary. • Output should be a clean CSV or Excel file containing at minimum two columns: Email and Phone. If the same record appears multiple times, the script must de-duplicate it automatically. • I would like the finished Python code (Scrapy, BeautifulSoup, Selenium, or any other proven library) so I can rerun it whenever the database grows. Please include a brief README explaining setup and usage. Acceptance criteria 1. A single data file containing all available email and phone entries from the site. 2. Script runs end-to-end on my machine without manual intervention beyond providing login or captcha keys if ...
...business directories and company websites to social platforms—and potentially any other page that reveals the information—I’m looking for a developer comfortable mixing and matching techniques: headless browsers for dynamic pages, straight HTTP requests where possible, and smart rate-limiting or proxy rotation to keep everything compliant and undetected. You’ll likely rely on Python (Scrapy, BeautifulSoup, Selenium or Playwright) as well as TypeScript where it makes sense, for example in a Node-based microservice or a lightweight dashboard that lets me trigger new scraping jobs and download the resulting datasets. I don’t mind which framework does what as long as the whole pipeline is clear, well-documented and easy for me to redeploy. Deliverables...
...Saudi-based firms with 50+ employees that show no internal audit positions in their organisational structure or job listings. • Record the company name, head-count shown on LinkedIn, industry, headquarters city, and a direct website or general contact email. I will review the file for accuracy, consistency of columns, and broken links before signing off. Feel free to work with Python, Selenium, BeautifulSoup, or the LinkedIn Sales Navigator API—whatever gives clean, export-ready results without violating platform terms....