i
GlobalLogic
Filter interviews by
I was interviewed in Dec 2024.
I applied via Approached by Company and was interviewed in Jun 2023. There were 3 interview rounds.
Spark internal flow involves job submission, DAG creation, task scheduling, and execution.
Job submission: User submits a Spark job to the SparkContext.
DAG creation: SparkContext creates a Directed Acyclic Graph (DAG) of the job.
Task scheduling: DAGScheduler breaks the DAG into stages and tasks, which are scheduled by TaskScheduler.
Task execution: Executors execute the tasks and return results to the driver.
Result fetch...
Hive meta store stores metadata about Hive tables, partitions, columns, and storage location.
Hive meta store is a central repository that stores metadata information about Hive tables, partitions, columns, and storage location.
It stores this metadata in a relational database like MySQL, Derby, or PostgreSQL.
The metadata includes information such as table names, column names, data types, file formats, and storage locati...
SQL query for simple tables
Use SELECT statement to retrieve data
Specify the columns you want to select
Use FROM clause to specify the tables you are querying from
Add WHERE clause to filter the results if needed
Top trending discussions
The aptitude test lasts 30 minutes and focuses on topics relevant to data engineering, including Spark, SQL, Azure, and PySpark.
The coding test is a one-hour examination on PySpark.
I applied via Naukri.com and was interviewed in Nov 2024. There was 1 interview round.
Enhanced optimization in AWS Glue improves job performance by automatically adjusting resources based on workload
Enhanced optimization in AWS Glue automatically adjusts resources like DPUs based on workload
It helps improve job performance by optimizing resource allocation
Users can enable enhanced optimization in AWS Glue job settings
Optimizing querying in Amazon Redshift involves proper table design, distribution keys, sort keys, and query optimization techniques.
Use appropriate distribution keys to evenly distribute data across nodes for parallel processing.
Utilize sort keys to physically order data on disk, reducing the need for sorting during queries.
Avoid using SELECT * and instead specify only the columns needed to reduce data transfer.
Use AN...
posted on 23 Dec 2024
I applied via Naukri.com and was interviewed in Jun 2024. There were 3 interview rounds.
Sample data and its transformations
Sample data can be in the form of CSV, JSON, or database tables
Transformations include cleaning, filtering, aggregating, and joining data
Examples: converting date formats, removing duplicates, calculating averages
Seeking new challenges and opportunities for growth in a more dynamic environment.
Looking for new challenges and opportunities for growth
Seeking a more dynamic work environment
Interested in expanding skill set and knowledge
Want to work on more innovative projects
I applied via Naukri.com and was interviewed in Sep 2024. There was 1 interview round.
SCD type 2 is a method used in data warehousing to track historical changes by creating a new record for each change.
SCD type 2 stands for Slowly Changing Dimension type 2
It involves creating a new record in the dimension table whenever there is a change in the data
The old record is marked as inactive and the new record is marked as current
It allows for historical tracking of changes in data over time
Example: If a cust...
posted on 4 Aug 2024
I am a Senior Data Engineer with 5+ years of experience in designing and implementing data pipelines for large-scale projects.
Experienced in ETL processes and data warehousing
Proficient in programming languages like Python, SQL, and Java
Skilled in working with big data technologies such as Hadoop, Spark, and Kafka
Strong understanding of data modeling and database management
Excellent problem-solving and communication sk
Developing a real-time data processing system for analyzing customer behavior on e-commerce platform.
Utilizing Apache Kafka for real-time data streaming
Implementing Spark for data processing and analysis
Creating machine learning models for customer segmentation
Integrating with Elasticsearch for data indexing and search functionality
I applied via Job Portal and was interviewed in Jul 2024. There was 1 interview round.
I have over 5 years of experience in data engineering, working with large datasets and implementing data pipelines.
Developed and maintained ETL processes to extract, transform, and load data from various sources
Optimized database performance and implemented data quality checks
Worked with cross-functional teams to design and implement data solutions
Utilized tools such as Apache Spark, Hadoop, and SQL for data processing
...
I would start by understanding the requirements, breaking down the task into smaller steps, researching if needed, and then creating a plan to execute the task efficiently.
Understand the requirements of the task
Break down the task into smaller steps
Research if needed to gather necessary information
Create a plan to execute the task efficiently
Communicate with stakeholders for clarification or updates
Regularly track prog
based on 3 interviews
Interview experience
based on 1 review
Rating in categories
Associate Analyst
3.9k
salaries
| ₹1 L/yr - ₹7.2 L/yr |
Senior Software Engineer
3.3k
salaries
| ₹5.3 L/yr - ₹22 L/yr |
Analyst
3.1k
salaries
| ₹1 L/yr - ₹5.5 L/yr |
Software Engineer
3k
salaries
| ₹3 L/yr - ₹13 L/yr |
Associate Consultant
2.8k
salaries
| ₹9.2 L/yr - ₹33.9 L/yr |
TCS
Wipro
Infosys
HCLTech