HomeAboutServicesPortfolioScrapersReviewsBlog Hire me
About

Data science,
without the distance.

datavyn is the data practice of Jamshaid Arif — Python developer and web-scraping specialist, working with web and business data since 2020.

What started as web scraping grew through data analysis into machine-learning engineering, and datavyn now covers the full journey: getting the data, understanding it, predicting with it, and presenting it so people actually use it.

The library and the studio

The work splits into two halves that feed each other. The first is the scraper library — more than 200 ready-to-use tools, published as Apify Actors, that pull structured data from e-commerce sites, real-estate portals, social platforms, job boards, business directories, and public data sources. Each tool has its own page describing exactly what it collects and returns, and every one delivers clean, deduplicated output in CSV or JSON. If a source you need isn't covered, a custom scraper is usually a short project away.

Every custom engagement sharpens the library, and the library accelerates every engagement. When a client needs competitor prices, an existing e-commerce scraper gets the first cut of data on day one while the custom pieces are built around it. When a new source gets scraped for a project, it often becomes a documented tool others can use later. That loop is why turnaround stays short without cutting corners.

From extraction to explanation

Collecting data is only the start. Most clients come with a question, not a URL: What will demand look like next quarter? Which listings are underpriced? Why did this metric move? Answering those takes the second half of the practice — exploratory analysis that finds the story in the data, forecasting models that project it forward, machine-learning techniques when the question is subtle, and dashboards that keep the answer alive as new data flows in.

The toolkit is Python end to end: pandas for wrangling; scikit-learn, TensorFlow, and PyTorch for modeling; PySpark when the data outgrows one machine; and modern scraping stacks — curl_cffi, httpx, BeautifulSoup, Playwright, Selenium — for extraction that holds up on real-world sites. The portfolio runs from recommender systems and predictive-maintenance models to retrieval-augmented clinical AI.

Who this is for

datavyn works best with founders, analysts, marketers, and researchers who need data work done properly but don't need a full-time data team. Typical engagements include price and competitor monitoring for e-commerce sellers, lead lists built from business directories, property-market datasets for investors, research datasets for academics, and forecasting or dashboard projects for small businesses that have data but no one to make sense of it.

Honest about the details

Two principles run through everything. First, the tools collect publicly available data only, and clients are responsible for using data lawfully — that's stated up front rather than discovered later. Second, no overpromising: scraping depends on what target sites expose, models are only as good as the history behind them, and a straight "that won't work, here's what will" is part of the service. You can read what clients say about that approach in 44 reviews on Fiverr (4.9★ average).

If you have a dataset that needs sense made of it, a website that needs turning into a spreadsheet, or a decision that needs numbers behind it, send a short note about what you're trying to do. You'll get a clear answer about feasibility, timeline, and cost — usually within two business days.

Technical skills

The stack behind the work.

Every technology listed here appears in shipped projects — the notebooks and Actors in the portfolio.

🐍

Core Python

Python · pandas · NumPy · Jupyter · Flask APIs · regular expressions · data cleaning & processing

🕸️

Scraping & Automation

curl_cffi · httpx · BeautifulSoup · Playwright · Selenium · anti-bot & TLS fingerprinting · proxy management · API integration

🤖

Machine Learning & AI

scikit-learn · TensorFlow / Keras · PyTorch · XGBoost · LightGBM · statsmodels · LangChain · ChromaDB · local LLMs (llama-cpp)

⚙️

Platforms & Scale

Apify platform & Actor development · PySpark for big data · Matplotlib / Seaborn visualization · CSV / JSON / Excel pipelines

Let's build something with data.

Describe your project and get a plain-English feasibility answer, timeline, and fixed quote.

Contact me →