• Characteristics of Big Data (5Vs)
• Traditional data vs Big Data
• Big Data use cases
• Big Data ecosystem overview
• Business benefits
• Hadoop architecture
• Hadoop components
• Hadoop distributions
• Hadoop use cases
• Hadoop advantages and limitations
• Master and slave nodes
• Hadoop daemons overview
• Single-node vs multi-node cluster
• Cluster setup basics
• High-level data flow
• NameNode and DataNode roles
• Block size and replication
• HDFS read and write process
• HDFS commands
• Fault tolerance in HDFS
• ResourceManager and NodeManager
• ApplicationMaster
• Resource allocation
• Scheduling concepts
• YARN benefits
• Mapper and Reducer concepts
• Data flow in MapReduce
• Writing MapReduce programs
• Job execution lifecycle
• MapReduce limitations
• Pig overview
• HBase basics
• Sqoop overview
• Flume overview
• Oozie overview
• Hive data types
• HiveQL basics
• Managed vs external tables
• Partitioning and bucketing
• Performance optimization
• Pig Latin language basics
• Data loading and storing
• Pig operators
• UDF concepts
• Pig performance considerations
• HBase data model
• Tables, rows, and column families
• HBase read and write operations
• HBase shell basics
• HBase use cases
• Flume for log data ingestion
• Kafka basics
• Real-time vs batch ingestion
• Data ingestion strategies
• Data pipeline design
• Oozie workflows
• Coordinators and bundles
• Scheduling Hadoop jobs
• Error handling
• Workflow monitoring
• Kerberos basics
• HDFS permissions
• YARN security
• Data encryption basics
• Security best practices







