임시 테이블 만들기

gilbird 2017. 1. 23. 09:07

2017. 1. 23. 09:07

참고: http://spark.apache.org/docs/1.6.2/sql-programming-guide.html

// sc is an existing SparkContext.
val sqlContext = new org.apache.spark.sql.SQLContext(sc)
// this is used to implicitly convert an RDD to a DataFrame.
import sqlContext.implicits._

// Define the schema using a case class.
// Note: Case classes in Scala 2.10 can support only up to 22 fields. To work around this limit,
// you can use custom classes that implement the Product interface.
case class Person(name: String, age: Int)

// Create an RDD of Person objects and register it as a table.
val people = sc.textFile("examples/src/main/resources/people.txt").map(_.split(",")).map(p => Person(p(0), p(1).trim.toInt)).toDF()
people.registerTempTable("people")

// SQL statements can be run by using the sql methods provided by sqlContext.
val teenagers = sqlContext.sql("SELECT name, age FROM people WHERE age >= 13 AND age <= 19")

// The results of SQL queries are DataFrames and support all the normal RDD operations.
// The columns of a row in the result can be accessed by field index:
teenagers.map(t => "Name: " + t(0)).collect().foreach(println)

// or by field name:
teenagers.map(t => "Name: " + t.getAs[String]("name")).collect().foreach(println)

// row.getValuesMap[T] retrieves multiple columns at once into a Map[String, T]
teenagers.map(_.getValuesMap[Any](List("name", "age"))).collect().foreach(println)
// Map("name" -> "Justin", "age" -> 19)

'Resources > Spark' 카테고리의 다른 글

spark-shell 사용법 (0)	2016.09.30
spark-shell에서 scala 버전 구하기 (0)	2016.09.29
spark-shell에서 s3 디렉토리를 지워야 하는 경우 external system command 사용하기 (0)	2016.02.13
Date에 Range를 넣어보자. (0)	2016.02.11

IT Lab

임시 테이블 만들기

'Resources > Spark' 카테고리의 다른 글

+ Recent posts

티스토리툴바