Skip to content
← Back to job listings

Data Scientist

gradera · Hyderabad, India Office

External listingfull-time13 days ago

About The Role

About GraderaGradera is an AI‑Native Services firm pioneering Software‑Orchestrated Services™—a new enterprise transformation model where software orchestrates human expertise, digital workers, and enterprise systems to deliver governed, scalable outcomes. We help enterprises move beyond fragmented AI pilots, disconnected automation, and labor‑led models by redesigning how work gets done across operations, product, engineering, customer experience, data, and core workflows.OverviewWe are seeking a highly analytical and curious Data Scientist to transform complex, real-world data into meaningful insights and scalable machine learning solutions. In this role, you will work across the full data lifecycle—partnering with data engineering and business teams to explore, clean, and understand diverse datasets, and translating those insights into models, experiments, and data-driven <recommendations.You> will play a critical role in bridging raw data and business impact, developing a deep understanding of how data is generated, structured, and used. This includes conducting rigorous exploratory analysis, assessing data quality and lineage, and building robust analytical datasets that power advanced modeling and reporting.This role offers the opportunity to work with large-scale data platforms, cloud infrastructure, and modern machine learning frameworks, while contributing to impactful decision-making through experimentation, analytics, and self-service data tools.Role & ResponsibilitiesCollect, clean, and analyze large structured and unstructured datasets from multiple internal and external sourcesConduct thorough exploratory data analysis (EDA) to understand data distributions, relationships, outliers, and missing value patternsProfile and audit datasets to assess data quality, completeness, consistency, and fitness for modelingInvestigate and document data lineage — understanding where data originates, how it flows, and how it transforms across systemsIdentify and resolve data anomalies, inconsistencies, and integrity issues in collaboration with data engineering teamsDevelop a deep understanding of the business domain and the underlying data that represents it — including what each field means, how it is captured, and what its limitations areTranslate raw, messy, real-world data into clean, well-understood analytical datasets ready for modeling and reportingApply statistical techniques such as correlation analysis, hypothesis testing, variance analysis, and distribution fitting to extract meaningful signals from noiseBuild and deploy machine learning models including regression, classification, clustering, NLP, and time-series analysisDesign, evaluate, and analyze A/B experiments and controlled tests using causal inference techniquesDevelop data-driven recommendations backed by rigorous statistical reasoningWrite clean, production-ready code in Python or RCollaborate with data engineers to build reliable data pipelines and feature storesDeploy and monitor ML models using MLOps best practices on cloud infrastructureBuild dashboards and self-serve analytics tools to support stakeholder decision-makingData Understanding & Analysis SkillsStrong ability to interrogate unfamiliar datasets and quickly develop a working understanding of their structure, semantics, and quirksExperience working with messy, incomplete, or poorly documented real-world dataSkilled in identifying hidden patterns, trends, seasonality, and anomalies through visual and statistical explorationAbility to ask the right questions about data — challenging assumptions, validating sources, and understanding the context in which data was collectedProficiency in data profiling, descriptive statistics, and summary reporting to communicate the shape and health of a datasetExperience creating data dictionaries, documentation, and data quality reports to support team-wide data understandingComfort working across structured (relational tables), semi-structured (JSON, XML), and unstructured (text, logs, sensor streams) data formatsTechnical Skills RequiredProficiency in Python (pandas, NumPy, scikit-learn, PyTorch or TensorFlow) and/or RStrong SQL skills with hands-on experience in DB2 and SQL ServerExperience with Databricks for large-scale data processing, feature engineering, and model trainingFamiliarity with cloud platforms: Azure or AWSExperience with data warehouses and big data platforms (Databricks, Snowflake, or Redshift)Knowledge of MLOps tools such as MLflow, Kubeflow, or AirflowExperience with streaming data technologies such as Kafka or SparkSolid foundation in probability, statistics, linear algebra, and experimental designNice to HaveExperience with deep learning, NLP, computer vision, or Bayesian methodsFamiliarity with real-time or streaming data pipelinesOpen-source contributions or published research

This is an external listing. JobSpring does not represent or verify the employer. Report this listing